← All projects

Jeve

A small economy that keeps running: four firms on one street and 424 simulated people. Every choice they make is a typed decision with a probability attached, not generated text.

September 2026

simulated people
424
simulated people
model cost of 5 simulated days
$0.04
model cost of 5 simulated days
tests across Python and the web app
419
tests across Python and the web app
decision records
62
decision records
Jeve
Python
Agent simulation
PostgreSQL
three.js
Next.js
FastAPI
Fly.io
Cloudflare

Most generative-agent simulations, like Stanford's Smallville or DeepMind's Concordia, run their agents on LLM prose. The agents talk to themselves in paragraphs and the world is built out of that text. It works, but it's expensive, it's hard to replay, and when something odd happens it's hard to tell why.

Jeve tests a different idea. Every choice an agent makes is one of three typed questions: a yes/no with a probability, a pick from a set of options, or a score on a scale. TypeSafe's Jev model answers them. Dialogue still exists, but only as something you can read. It's rendered when you click, and no code that runs the world is allowed to read it back.

The voxel town with one employee selected. The side panel shows her temperament and the probabilities Jev gave for her last decisions.
Clicking a person shows their temperament and the full distribution behind each choice they just made.

The street

Four businesses share one block, and they depend on each other:

Firm What it does What it depends on
Tallybird Software Sells the time tracking, invoicing and point-of-sale software the others use Subscriptions
Halloran & Pike LLP A law firm that bills by the hour Tallybird's time tracking and invoicing
Ledgerline Accounting Payroll and the monthly close for the street Invoicing, and clients sending data on time
Third Rail Cafe Feeds everyone Tallybird's point-of-sale and time tracking

Ten kinds of typed flows run over a double-entry ledger kept in integer cents, including invoices, support tickets, incidents, payroll, credit, catering and lunch. Nobody scripts the drama. If Tallybird's invoicing goes down at the end of the month, bills go out late, payments come in late, someone's cash runs low, and that firm starts paying its own suppliers late.

Close-up of the Tallybird Software office with the firm's books in the side panel.
Clicking a building opens its books. Here Tallybird is in the middle of an invoicing outage with 106 unpaid invoices.

How it works

The simulation is about 23,000 lines of Python on a simulated clock that ticks in 15-minute steps. Each tick is a single Postgres transaction that commits the world changes, the events, the decisions and the clock advance together. Every event records the IDs of the events that caused it, so the frontend can show a causal timeline instead of a chat log.

Decisions are batched. Agents in the same situation share one model call, and only the random draw is per person, which keeps the whole thing cheap. Five simulated days produced 2,082 events and 5,823 decisions and cost $0.0436 in model calls.

The spending limit controls the clock speed. When the daily budget runs low, the simulation slows down before it stops. If the model is slow, the world waits for it. It never fills in a default answer.

The site is a static Next.js export with a three.js voxel town. The API and the simulation daemon run as two processes of one app on Fly.io.

The Jeve landing page: the voxel town above a dashboard of each firm's cash, receivables and recent decisions.
The landing page. The town is live, and the dashboard below it lists each person's decisions with the random draw that picked each one.
The same street at dusk and at 2 AM.
The town follows the simulated clock, from evening to 2 AM.

The hard parts

Every run replays exactly

A random number generator carries state that you then have to save. Jeve avoids that entirely. Every random draw is a hash of the seed, the agent, the decision number and the question. With recorded model responses, a whole run replays byte for byte, and the full end-to-end test runs offline for $0.00.

It survives being killed

I tested crash recovery by killing the daemon with kill -9 in the middle of a tick. When it restarts, the event log has no gaps and no duplicates, and the ledger still balances.

The architecture is enforced by a test

The Python code is split into layers. A test walks the import graph and fails if any code that computes the world can reach generated text, including imports hidden inside functions. The rule that "prose is only a view" can't erode quietly.

Checking the risky assumption early

The whole idea fails if the model ignores personality. Before building the world, I ran a probe: 280 live calls that changed only an agent's temperament. All 6 test cases produced different distributions, and the control case stayed flat. It cost less than a cent.

Measuring instead of guessing

Meetings where something is at stake can run up to three rounds instead of being decided in one shot. To see whether that mattered, I played the same 21 simulated days twice across five seeds, once each way, using the cheaper rules-based version of the agents. With rounds, outage minutes dropped from 2,703 to 2,058 while decisions rose 2.4%. A separate 35-day soak run held all 12 invariants I check for.

Where it stands

Jeve runs in production and publishes field reports from inside the world, such as how long a failing firm lasted. I kept 62 decision records along the way. Some features are still partial. A second-opinion model for uncertain, high-stakes calls runs in shadow mode only, and its answers don't change the world yet.