Most generative-agent simulations, like Stanford's Smallville or DeepMind's Concordia, run their agents on LLM prose. The agents talk to themselves in paragraphs and the world is built out of that text. It works, but it's expensive, it's hard to replay, and when something odd happens it's hard to tell why.
Jeve tests a different idea. Every choice an agent makes is one of three typed questions: a yes/no with a probability, a pick from a set of options, or a score on a scale. TypeSafe's Jev model answers them. Dialogue still exists, but only as something you can read. It's rendered when you click, and no code that runs the world is allowed to read it back.

The street
Four businesses share one block, and they depend on each other:
| Firm | What it does | What it depends on |
|---|---|---|
| Tallybird Software | Sells the time tracking, invoicing and point-of-sale software the others use | Subscriptions |
| Halloran & Pike LLP | A law firm that bills by the hour | Tallybird's time tracking and invoicing |
| Ledgerline Accounting | Payroll and the monthly close for the street | Invoicing, and clients sending data on time |
| Third Rail Cafe | Feeds everyone | Tallybird's point-of-sale and time tracking |
Ten kinds of typed flows run over a double-entry ledger kept in integer cents, including invoices, support tickets, incidents, payroll, credit, catering and lunch. Nobody scripts the drama. If Tallybird's invoicing goes down at the end of the month, bills go out late, payments come in late, someone's cash runs low, and that firm starts paying its own suppliers late.

How it works
The simulation is about 23,000 lines of Python on a simulated clock that ticks in 15-minute steps. Each tick is a single Postgres transaction that commits the world changes, the events, the decisions and the clock advance together. Every event records the IDs of the events that caused it, so the frontend can show a causal timeline instead of a chat log.
Decisions are batched. Agents in the same situation share one model call, and only the random draw is per person, which keeps the whole thing cheap. Five simulated days produced 2,082 events and 5,823 decisions and cost $0.0436 in model calls.
The spending limit controls the clock speed. When the daily budget runs low, the simulation slows down before it stops. If the model is slow, the world waits for it. It never fills in a default answer.
The site is a static Next.js export with a three.js voxel town. The API and the simulation daemon run as two processes of one app on Fly.io.


The hard parts
Every run replays exactly
A random number generator carries state that you then have to save. Jeve avoids that entirely. Every random draw is a hash of the seed, the agent, the decision number and the question. With recorded model responses, a whole run replays byte for byte, and the full end-to-end test runs offline for $0.00.
It survives being killed
I tested crash recovery by killing the daemon with kill -9 in the middle of a tick. When it restarts, the event log has no gaps and no duplicates, and the ledger still balances.
The architecture is enforced by a test
The Python code is split into layers. A test walks the import graph and fails if any code that computes the world can reach generated text, including imports hidden inside functions. The rule that "prose is only a view" can't erode quietly.
Checking the risky assumption early
The whole idea fails if the model ignores personality. Before building the world, I ran a probe: 280 live calls that changed only an agent's temperament. All 6 test cases produced different distributions, and the control case stayed flat. It cost less than a cent.
Measuring instead of guessing
Meetings where something is at stake can run up to three rounds instead of being decided in one shot. To see whether that mattered, I played the same 21 simulated days twice across five seeds, once each way, using the cheaper rules-based version of the agents. With rounds, outage minutes dropped from 2,703 to 2,058 while decisions rose 2.4%. A separate 35-day soak run held all 12 invariants I check for.
Where it stands
Jeve runs in production and publishes field reports from inside the world, such as how long a failing firm lasted. I kept 62 decision records along the way. Some features are still partial. A second-opinion model for uncertain, high-stakes calls runs in shadow mode only, and its answers don't change the world yet.
