← All projects

Workbench

Reinforcement learning environments built from realistic office work. A multi-agent simulation runs a law firm for months, and AI agents are graded on tasks whose answers come from that recorded world.

August 2026

commits in 23 days
913
commits in 23 days
emulated tools across 7 servers
85
emulated tools across 7 servers
certified tasks
8
certified tasks
test functions
1,538
test functions
Python
Reinforcement learning
LLM evaluation
Multi-agent simulation
MCP
DSPy
SQLite
Docker

Many AI agent benchmarks are puzzles in disguise. Mechanical tasks are easy for frontier models. In one of my own earlier simulated firms, a frontier model scored about 1.000 on every task no matter how much harder I made them. That doesn't tell you whether a model can do a real job.

Workbench builds the job instead. It simulates a whole professional firm for months, then drops an agent into that firm's inbox, calendar, chat and document system and asks it to do the kind of work an administrator or associate would do. The goal is tasks where scores stay between roughly 0.2 and 0.8 across three model tiers. Just as important is proving that a low score is the model's mistake, not a bug in the environment, the answer key or the harness.

The Workbench pipeline: workplace spec, multi-agent simulation, world log, environment, and graded tasks.
The pipeline from a description of a firm to graded agent runs.

How it works

It starts with a description of a firm: the staff, the client list, the communication channels, seed documents and a season of client demand. A simulation engine then plays every employee and client, day by day. Each morning people plan their day around their real calendar. They wake on a 30-minute grid, check their memory, and act. Meetings produce turn-by-turn transcripts. Client requests arrive on a schedule, and at the end of each day people reflect and fill in timesheets.

The engine follows the pattern DeepMind's Concordia introduced, but I rebuilt it from scratch without using any of its code. A "game master" that makes no model calls checks every action against the state of the world and rejects anything that refers to something that doesn't exist.

The output is a typed, append-only world log. From that log, Workbench builds an environment: the firm's documents plus emulated versions of Gmail, Google Calendar, Slack, iManage and Clio, each served as an MCP tool backed by SQLite. The agent never sees the simulation. It only has the documents and the tools, like a new hire would.

A task, for example

The main task pack is a 31-person law firm, Merrick Stanton LLP, simulated over 180 days with 567 meetings. One task starts like this:

You are the practice administrator at Merrick Stanton LLP, a litigation-and-transactions firm. The firm's standing meetings are where dates get set. Somebody says they will have a thing done by a day; a week later, in the same standing meeting, they say a different day.

The agent has to produce a register of what is still owed: who owes it, in which meeting they said so, the date it's due, and how many earlier promises it replaced. None of that is written down anywhere except in six months of transcripts. Deadlines are relative, too. "EOD" said in January and "EOD" said in June are five months apart, so the task grades the resolved date. An agent that never opens a transcript scores nothing, and one that takes each person's first statement at face value scores almost nothing.

Model tier Trials Mean score
opus-5 0.643, 0.704, 0.769 0.706
glm-5.2 0.454, 0.500, 0.530 0.495
kimi-k3 0.428, 0.467, 0.581 0.492

These are the scores recorded in the repo for that task. An empty register scores 0.200 and reporting every candidate scores 0.000, so the numbers above mean something.

The hard parts

Replaying an LLM simulation byte for byte

Every model call is stored under a hash of its exact request, so a recorded world replays offline with no network. Getting the same log regardless of how many agents run at once took more work. The engine only admits events in a fixed canonical order, memory retrieval uses integer scoring so floating point can't drift, and the run database commits after every step so a killed run resumes where it stopped. CI replays a recorded world three ways, sequentially, in parallel, and killed then resumed, and checks that all three match.

The lint config bans random, uuid4, time.time and datetime.now outright.

Answer keys I can defend

Every grading rule is derived twice, independently. One version matches character spans and the other matches word tokens. Across 4,998 meeting turns they agree with zero disagreements. Other scripts find rule conditions that never fire and check that every answer is reachable through the real tools. Disputed rows get written adjudications.

Testing the tests

The quality gates are mutation-tested in CI. A script breaks each safety check on purpose in 37 different ways and fails the build if any of those breaks goes unnoticed.

Terminal output from the Workbench test suite, an answer key summary and the task index.
Real output: the vendor parity tests, the shape of one task's answer key, and the task index.

Tools that match the real products

Each emulated product is checked against a dated snapshot of the vendor's official tool list. CI fails if a tool is missing, if one was invented, or if required parameters don't match. Inside the container, a small setuid C program lets the tool servers read databases that the agent's own user can't open.

Honest reporting

The certification script refuses to count a run that didn't finish as a zero. It reports completion rates next to scores instead. I also kept a list of 50 "measured laws", things I learned by measuring, each with the number and what it cost to find out.

Where it stands

The shipping pack has 12 live tasks, 8 of them certified against three model tiers, with the command to reproduce each score written down. I also packaged the method as a Claude Code plugin with 8 skills, so the same approach can be used to build environments for other kinds of work.