Drop anything into a town of AI people who live for days, remember, talk to each other and react together.
Python
0
4 commits
updated Sep 30, 2026
Drop anything into a town of 200 AI people who live for days, remember, talk to each other and react together.
Populace turns a sentence - "a commuter suburb of 200 people with a high street, a station, a grocery, a diner, a pharmacy and a school" - into a town of residents, each with a home, a job or a school, a household, money, needs, a schedule and people they know. It runs them headless for as many in-game days as you like, lets you drop things in while it runs (a new shop, a price rise, an outage, a letter, a notice, a newcomer, or your own AI agent that residents can text, phone or visit), and writes a plain-language report of what the town did about it.
A resident knows only what they saw, heard or were told. News travels by people being in the same place and talking. Nobody is instructed to react to what you drop in; they decide for themselves, one model call at a time.
MIT licensed. Formerly called worldsim; built from the brain of Alive, a life-sim where every character is an LLM agent.
What residents said to their internet provider, word for word, in the first four-day run (Qwen3-32B; the appendix of its report has every transcript):
Luis Vargas, on the phone to the internet provider, Day 1 18:00: "Why did my bill jump to $184.60 this month?"
Ximena Aguilar, by text, Day 1 18:30: "I'm Ximena Aguilar. I'm with Victoria Court. My line's been down since 07:00. Rose Grant already raised a fault with you, but I just wanted to check in myself."
Will Ward, whose internet at home was fine, Day 3 16:00: "It's Will Ward. The internet's still down. It was supposed to be back by tomorrow."
Paul Shaw, to Edward Shaw, Day 4 17:00: "I've had my share of trouble with Northline. Best to get on to them quick."
The packaged demo (populace demo isp) puts an internet provider's helpdesk,
"Northline", in front of a 200-person suburb for four in-game days and
schedules trouble: the internet off on one street, twelve wrong bills by text,
slow internet in 25 homes, and the first street off again on the third
evening. We ran it twice on the same town, seed, schedule and server, with
the residents driven by Qwen3-32B (Q4_K_M) on one RTX 4090. Only the helpdesk
changed:
examples/llm_helpdesk.py): the same
Qwen3-32B with a support agent's prompt, the customer's account on screen
(plan, bill, live line status, the engineers' diary) and four actions:
correct a bill, credit, send an engineer, promise.populace/demos/isp.py):
keyword matching and a ticket per customer, a deliberately simple baseline.From each run's computed report
(LLM, run live-32b-4d-llm;
rules, run live-32b-4d-rules):
| LLM helpdesk | Rule-based bot | |
|---|---|---|
| Residents with a problem the helpdesk handles | 100 | 100 |
| ...who tried to get in touch / reached it | 39 / 25 | 48 / 34 |
| Contacts answered / all | 33 / 57 | 41 / 65 |
| Problems the helpdesk fixed on Day 1 / in all | 7 / 16 | 4 / 31 |
| Contacts about a problem nobody at home had, and acted on | 8, acted on 3 | 10, acted on 4 |
| Promises made | 9 | 31 |
| ...kept / broken / not yet due at the end | 2 / 6 / 1 | 28 / 2 / 1 |
| Still broken at the end | 8 | 4 |
| Minutes per in-game day | 10.1 | 10.8 |
Where the language model was better:
Where it was worse:
Neither is the good agent. The point is that one run of a town shows where one agent beats another, in residents' own words, with every contact transcribed in the report. Plug in your own - any Python object or HTTP endpoint (docs/AGENTS.md) - and run it against the same town:
.venv/bin/populace demo isp --agent northline=path/to/your_agent.py --model-url http://127.0.0.1:8080/v1
The first live run of the demo, on earlier engine code, with the rule-based
bot answering. It is there to show the town finding faults in a support
agent, not as an example of a good one. From the run's computed report
(docs/samples/isp-200-live-32b-report.md,
run live-32b-4d):
| Residents with a problem the helpdesk handles | 100 |
| ...who tried to get in touch | 40 |
| ...who reached it | 27 (13 only ever got the out-of-hours recording) |
| Contacts | 59: 58 texts, 1 phone call; 35 answered, 24 out of hours |
| Median hours from a problem starting to getting in touch about it | 7.2 |
| Promises the helpdesk made / kept / broken / not yet due | 26 / 21 / 2 / 3 (both "broken" were judged by an older rule, since fixed, that let a later outage break them) |
| Problems the helpdesk fixed / that ended on the schedule anyway | 39 / 100 |
| Still broken at the end | 6 |
What the town found wrong with the helpdesk:
The report keeps these apart from the town's own mistakes (below), in two labelled groups: What the agent got wrong and Where the simulation is weak.
live-32b-4d); about 55 to 58 minutes per
day for 50 residents with the 7B on an M1 Mac. Calls per day are set by the
preset, not the population, so a bigger town costs about the same per day.repeated_line 3,
echo 3 in this run).Python 3.11 or newer.
git clone https://github.com/populace-sim/populace && cd populace
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest -q # 294 tests, no model needed
.venv/bin/populace demo isp # the ISP demo in mock: 200 people, 4 days, under a minute
The report is written next to the run's logs, and its path is printed.
On Windows the virtualenv keeps its programs in .venv\Scripts\ rather
than .venv/bin/, and python3 is usually py:
py -m venv .venv
.venv\Scripts\pip install -e ".[dev]"
.venv\Scripts\python -m pytest -q
.venv\Scripts\populace demo isp
Everywhere below, read .venv/bin/X as .venv\Scripts\X. Paths given to
Populace itself (towns/brookhaven, examples/llm_helpdesk.py) work with
either slash.
Live, with a real model - any OpenAI-compatible server. On a 24 GB GPU with llama.cpp:
llama-server -m Qwen3-32B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 -c 18432 -np 2 --cache-reuse 256 --jinja --reasoning-budget 0 \
--chat-template-kwargs '{"enable_thinking":false}'
.venv/bin/populace demo isp --model-url http://127.0.0.1:8080/v1 --concurrency 2 --profile compact
No GPU? Use a hosted API. Any OpenAI-compatible endpoint works, for
example OpenRouter serving Qwen3-32B. The key is read from POPULACE_API_KEY
only and is never written to a config file, a log, a report or a run folder.
Turn the model's thinking off per request: --thinking-off openrouter uses
OpenRouter's reasoning setting, and --thinking-off no-think appends Qwen3's
/no_think to each prompt for hosts with no such setting. Other hosts may name
models and switches differently, so check their docs.
export POPULACE_API_KEY=your-key-here
.venv/bin/populace demo isp --model-url https://openrouter.ai/api/v1 --model qwen/qwen3-32b \
--thinking-off openrouter --concurrency 8 --profile compact
Two cautions. A hosted run costs money, and Populace has no built-in spending cap: four days of the demo is roughly 800 model calls. Every run prints its model calls and tokens in and out at the end, and the report lists them, so you can see what it used. A hosted Qwen3-32B may behave differently from the published 4090 runs (a different quantisation, serving stack or chat template), so compare like with like.
Your own town:
.venv/bin/populace new "a small town of 50 people with a diner, a grocery and a workshop" --seed 3
.venv/bin/populace inject towns/brookhaven --file examples/injections.json --write
.venv/bin/populace run towns/brookhaven --days 2 --agent northline=helpdesk
.venv/bin/populace report towns/brookhaven/runs/<run_id>
examples/injections.json is written for this town (the seed-3 Brookhaven):
it registers a helpdesk, turns the internet off on Station Crescent, shuts the
diner and posts a notice. Its last line is refused on purpose, to show a
refusal: a notice saying "Everyone knows..." reaches into people's minds.
Without --write the inject command only checks.
handle(message, ctx), or any
HTTP server. It sees what a real service would know about a customer, and
changes the town only through what it does: fix, credit, send someone
round, promise. docs/AGENTS.md.
examples/llm_helpdesk.py is a model-backed
helpdesk that shares the residents' model server:
populace demo isp --agent northline=examples/llm_helpdesk.py (free in
mock; add --model-url for live).Each half-hour tick, a scheduler spends a fixed budget of model calls on the residents with the best reason to think - someone spoke to them, they saw something, their phone went, they are hungry, or they have not thought for a while - and everybody else carries on. Conversations are one speaker per call, so nobody's private state leaks into anyone else's words. Names are learned only by being said aloud; the model knows residents by opaque ids. Each night, people reflect on their day. Every failure an action can meet is said to the resident out loud; nothing fails silently.
docs/STATUS.md is the build log: what is verified, the decisions behind it, and what is open.
MIT licensed.
Python
100.0%
Drop anything into a town of AI people who live for days, remember, talk to each other and react together.
Python
0
4 commits
updated Sep 30, 2026
Drop anything into a town of 200 AI people who live for days, remember, talk to each other and react together.
Populace turns a sentence - "a commuter suburb of 200 people with a high street, a station, a grocery, a diner, a pharmacy and a school" - into a town of residents, each with a home, a job or a school, a household, money, needs, a schedule and people they know. It runs them headless for as many in-game days as you like, lets you drop things in while it runs (a new shop, a price rise, an outage, a letter, a notice, a newcomer, or your own AI agent that residents can text, phone or visit), and writes a plain-language report of what the town did about it.
A resident knows only what they saw, heard or were told. News travels by people being in the same place and talking. Nobody is instructed to react to what you drop in; they decide for themselves, one model call at a time.
MIT licensed. Formerly called worldsim; built from the brain of Alive, a life-sim where every character is an LLM agent.
What residents said to their internet provider, word for word, in the first four-day run (Qwen3-32B; the appendix of its report has every transcript):
Luis Vargas, on the phone to the internet provider, Day 1 18:00: "Why did my bill jump to $184.60 this month?"
Ximena Aguilar, by text, Day 1 18:30: "I'm Ximena Aguilar. I'm with Victoria Court. My line's been down since 07:00. Rose Grant already raised a fault with you, but I just wanted to check in myself."
Will Ward, whose internet at home was fine, Day 3 16:00: "It's Will Ward. The internet's still down. It was supposed to be back by tomorrow."
Paul Shaw, to Edward Shaw, Day 4 17:00: "I've had my share of trouble with Northline. Best to get on to them quick."
The packaged demo (populace demo isp) puts an internet provider's helpdesk,
"Northline", in front of a 200-person suburb for four in-game days and
schedules trouble: the internet off on one street, twelve wrong bills by text,
slow internet in 25 homes, and the first street off again on the third
evening. We ran it twice on the same town, seed, schedule and server, with
the residents driven by Qwen3-32B (Q4_K_M) on one RTX 4090. Only the helpdesk
changed:
examples/llm_helpdesk.py): the same
Qwen3-32B with a support agent's prompt, the customer's account on screen
(plan, bill, live line status, the engineers' diary) and four actions:
correct a bill, credit, send an engineer, promise.populace/demos/isp.py):
keyword matching and a ticket per customer, a deliberately simple baseline.From each run's computed report
(LLM, run live-32b-4d-llm;
rules, run live-32b-4d-rules):
| LLM helpdesk | Rule-based bot | |
|---|---|---|
| Residents with a problem the helpdesk handles | 100 | 100 |
| ...who tried to get in touch / reached it | 39 / 25 | 48 / 34 |
| Contacts answered / all | 33 / 57 | 41 / 65 |
| Problems the helpdesk fixed on Day 1 / in all | 7 / 16 | 4 / 31 |
| Contacts about a problem nobody at home had, and acted on | 8, acted on 3 | 10, acted on 4 |
| Promises made | 9 | 31 |
| ...kept / broken / not yet due at the end | 2 / 6 / 1 | 28 / 2 / 1 |
| Still broken at the end | 8 | 4 |
| Minutes per in-game day | 10.1 | 10.8 |
Where the language model was better:
Where it was worse:
Neither is the good agent. The point is that one run of a town shows where one agent beats another, in residents' own words, with every contact transcribed in the report. Plug in your own - any Python object or HTTP endpoint (docs/AGENTS.md) - and run it against the same town:
.venv/bin/populace demo isp --agent northline=path/to/your_agent.py --model-url http://127.0.0.1:8080/v1
The first live run of the demo, on earlier engine code, with the rule-based
bot answering. It is there to show the town finding faults in a support
agent, not as an example of a good one. From the run's computed report
(docs/samples/isp-200-live-32b-report.md,
run live-32b-4d):
| Residents with a problem the helpdesk handles | 100 |
| ...who tried to get in touch | 40 |
| ...who reached it | 27 (13 only ever got the out-of-hours recording) |
| Contacts | 59: 58 texts, 1 phone call; 35 answered, 24 out of hours |
| Median hours from a problem starting to getting in touch about it | 7.2 |
| Promises the helpdesk made / kept / broken / not yet due | 26 / 21 / 2 / 3 (both "broken" were judged by an older rule, since fixed, that let a later outage break them) |
| Problems the helpdesk fixed / that ended on the schedule anyway | 39 / 100 |
| Still broken at the end | 6 |
What the town found wrong with the helpdesk:
The report keeps these apart from the town's own mistakes (below), in two labelled groups: What the agent got wrong and Where the simulation is weak.
live-32b-4d); about 55 to 58 minutes per
day for 50 residents with the 7B on an M1 Mac. Calls per day are set by the
preset, not the population, so a bigger town costs about the same per day.repeated_line 3,
echo 3 in this run).Python 3.11 or newer.
git clone https://github.com/populace-sim/populace && cd populace
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest -q # 294 tests, no model needed
.venv/bin/populace demo isp # the ISP demo in mock: 200 people, 4 days, under a minute
The report is written next to the run's logs, and its path is printed.
On Windows the virtualenv keeps its programs in .venv\Scripts\ rather
than .venv/bin/, and python3 is usually py:
py -m venv .venv
.venv\Scripts\pip install -e ".[dev]"
.venv\Scripts\python -m pytest -q
.venv\Scripts\populace demo isp
Everywhere below, read .venv/bin/X as .venv\Scripts\X. Paths given to
Populace itself (towns/brookhaven, examples/llm_helpdesk.py) work with
either slash.
Live, with a real model - any OpenAI-compatible server. On a 24 GB GPU with llama.cpp:
llama-server -m Qwen3-32B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 -c 18432 -np 2 --cache-reuse 256 --jinja --reasoning-budget 0 \
--chat-template-kwargs '{"enable_thinking":false}'
.venv/bin/populace demo isp --model-url http://127.0.0.1:8080/v1 --concurrency 2 --profile compact
No GPU? Use a hosted API. Any OpenAI-compatible endpoint works, for
example OpenRouter serving Qwen3-32B. The key is read from POPULACE_API_KEY
only and is never written to a config file, a log, a report or a run folder.
Turn the model's thinking off per request: --thinking-off openrouter uses
OpenRouter's reasoning setting, and --thinking-off no-think appends Qwen3's
/no_think to each prompt for hosts with no such setting. Other hosts may name
models and switches differently, so check their docs.
export POPULACE_API_KEY=your-key-here
.venv/bin/populace demo isp --model-url https://openrouter.ai/api/v1 --model qwen/qwen3-32b \
--thinking-off openrouter --concurrency 8 --profile compact
Two cautions. A hosted run costs money, and Populace has no built-in spending cap: four days of the demo is roughly 800 model calls. Every run prints its model calls and tokens in and out at the end, and the report lists them, so you can see what it used. A hosted Qwen3-32B may behave differently from the published 4090 runs (a different quantisation, serving stack or chat template), so compare like with like.
Your own town:
.venv/bin/populace new "a small town of 50 people with a diner, a grocery and a workshop" --seed 3
.venv/bin/populace inject towns/brookhaven --file examples/injections.json --write
.venv/bin/populace run towns/brookhaven --days 2 --agent northline=helpdesk
.venv/bin/populace report towns/brookhaven/runs/<run_id>
examples/injections.json is written for this town (the seed-3 Brookhaven):
it registers a helpdesk, turns the internet off on Station Crescent, shuts the
diner and posts a notice. Its last line is refused on purpose, to show a
refusal: a notice saying "Everyone knows..." reaches into people's minds.
Without --write the inject command only checks.
handle(message, ctx), or any
HTTP server. It sees what a real service would know about a customer, and
changes the town only through what it does: fix, credit, send someone
round, promise. docs/AGENTS.md.
examples/llm_helpdesk.py is a model-backed
helpdesk that shares the residents' model server:
populace demo isp --agent northline=examples/llm_helpdesk.py (free in
mock; add --model-url for live).Each half-hour tick, a scheduler spends a fixed budget of model calls on the residents with the best reason to think - someone spoke to them, they saw something, their phone went, they are hungry, or they have not thought for a while - and everybody else carries on. Conversations are one speaker per call, so nobody's private state leaks into anyone else's words. Names are learned only by being said aloud; the model knows residents by opaque ids. Each night, people reflect on their day. Every failure an action can meet is said to the resident out loud; nothing fails silently.
docs/STATUS.md is the build log: what is verified, the decisions behind it, and what is open.
MIT licensed.
Python
100.0%