← All projects

Mimic

Answer about 30 questions and Mimic learns how you make decisions. It guesses each answer before you give it, then shows you whether it was right.

September 2026

test cases
473
test cases
predictors compared head to head
6
predictors compared head to head
replayed states matched exactly
32 / 32
replayed states matched exactly
model cost of a 30-question session
$0.09
model cost of a 30-question session
TypeScript
Next.js
Cloudflare Workers
LLM evaluation
Active learning
D1
Vectorize
Queues

A lot of "digital twin" projects ask an LLM to pretend to be you. They are hard to trust. Many barely beat a guess made from your demographics, and they are usually graded on the same answers they learned from, which makes them look better than they are.

Mimic is my attempt to measure this honestly. You answer short, concrete questions about how you'd act. Before each answer, Mimic commits to a prediction. Then it shows you what it guessed, and it keeps score in a way that can't flatter itself.

A Mimic session. The left panel tracks how well the mimic knows you; the right shows its guess next to your answer.
Question 14 of about 30. The mimic guessed wrong here, and the panel on the left updates to show what changed.

How it works

Every question is a real decision with a fixed set of answers: pick one option, say yes or no, or rate yourself on a scale. The mimic returns a probability for each answer, you answer, and then the guess is revealed. It says "Missed" when it was wrong and "Close" when a scale answer was near.

Three phone screens: a question, the reveal after answering, and a scale question with the predicted distribution.
On a phone: the question, the reveal, and a scale question with the mimic's full distribution.

The score it shows you is called fidelity. It divides the mimic's accuracy by your own consistency, which Mimic measures by occasionally asking you something again. If you only agree with yourself 80% of the time, a mimic that matches you 80% of the time knows you as well as you know yourself.

It also picks questions deliberately. It keeps a running estimate of what it's unsure about for you and asks the question that should teach it the most. Standard strategies like entropy and BALD run alongside it as controls so the choice can be measured.

A scale question showing the mimic's predicted distribution over five answers.
For scale questions the mimic predicts a full distribution, not a single number.

You own the result. You can download your profile as a file to hand to your own agents, and deleting your mimic removes it from every store. Sensitive topics such as politics, religion and health are off until you turn them on, and those parts of the model only learn from answers you gave directly.

The topics and consent sheet with opt-in toggles for sensitive areas.
Sensitive areas are opt-in, and you can change them mid-session.

Under the hood

The app is a Next.js 16 site running on Cloudflare Workers. A separate Worker handles background jobs from queues: generating questions, running the other predictors, and updating what the mimic has learned. Data lives in D1, R2, KV and Vectorize. The prediction logic is a plain TypeScript package with no Cloudflare or Node imports, so the same code runs in production, in tests and in the offline eval tools.

The main predictor is Jev, a decision model from TypeSafe. Five inexpensive LLMs make "shadow" predictions on the same questions, from the same saved state, so they can be compared fairly.

The hard parts

Predictions are sealed before you answer

This is the rule the whole project rests on. When Mimic picks the next question, it computes and stores its prediction, along with a hash of everything it knew at that moment, before the question reaches your screen. Scoring only ever uses those stored predictions. A baseline that sees only your profile runs on every question too, so there is always something to compare against.

Because every prediction records exactly what it was based on, old states can be rebuilt. A replay test rebuilt 32 of 32 saved states with matching hashes. It also caught a real bug where learned traits were leaking into states they shouldn't have been in.

A lab for comparing predictors

The /lab dashboard puts all six predictors side by side on accuracy, log loss, Brier score, calibration error, cost and latency. It also watches the rules above. If a prediction was ever shown before it was sealed, the dashboard would flag it.

The lab dashboard comparing predictors on accuracy, log loss, calibration, cost and latency.
The lab page compares every predictor on identical sealed states. These numbers are from test runs, not real people.

Honest prompt optimization

I built an evaluation and prompt-optimization harness (GEPA) that runs from GitHub Actions, with spend caps, a check for answer leakage, and a held-out set. The first time it found an improvement in production, the gain didn't hold up on held-out data. So I changed the rule: the harness now reports an improvement only when it replicates.

Privacy that is tested

Tests try to infer a person's religion from answers they didn't consent to share, and they fail if any sensitive trait gets filled in that way. Another test checks that deleting a mimic really clears it from every store.

Where it stands

Mimic is deployed and invite-only. The first production report covers 4 people who opted in and 265 questions, which is too few to claim anything about accuracy yet, and the README says so. What's done is the part that makes a later claim believable: sealed predictions, a baseline on every question, replayable state and a lab that shows the numbers.