A lot of "digital twin" projects ask an LLM to pretend to be you. They are hard to trust. Many barely beat a guess made from your demographics, and they are usually graded on the same answers they learned from, which makes them look better than they are.
Mimic is my attempt to measure this honestly. You answer short, concrete questions about how you'd act. Before each answer, Mimic commits to a prediction. Then it shows you what it guessed, and it keeps score in a way that can't flatter itself.

How it works
Every question is a real decision with a fixed set of answers: pick one option, say yes or no, or rate yourself on a scale. The mimic returns a probability for each answer, you answer, and then the guess is revealed. It says "Missed" when it was wrong and "Close" when a scale answer was near.

The score it shows you is called fidelity. It divides the mimic's accuracy by your own consistency, which Mimic measures by occasionally asking you something again. If you only agree with yourself 80% of the time, a mimic that matches you 80% of the time knows you as well as you know yourself.
It also picks questions deliberately. It keeps a running estimate of what it's unsure about for you and asks the question that should teach it the most. Standard strategies like entropy and BALD run alongside it as controls so the choice can be measured.

You own the result. You can download your profile as a file to hand to your own agents, and deleting your mimic removes it from every store. Sensitive topics such as politics, religion and health are off until you turn them on, and those parts of the model only learn from answers you gave directly.

Under the hood
The app is a Next.js 16 site running on Cloudflare Workers. A separate Worker handles background jobs from queues: generating questions, running the other predictors, and updating what the mimic has learned. Data lives in D1, R2, KV and Vectorize. The prediction logic is a plain TypeScript package with no Cloudflare or Node imports, so the same code runs in production, in tests and in the offline eval tools.
The main predictor is Jev, a decision model from TypeSafe. Five inexpensive LLMs make "shadow" predictions on the same questions, from the same saved state, so they can be compared fairly.
The hard parts
Predictions are sealed before you answer
This is the rule the whole project rests on. When Mimic picks the next question, it computes and stores its prediction, along with a hash of everything it knew at that moment, before the question reaches your screen. Scoring only ever uses those stored predictions. A baseline that sees only your profile runs on every question too, so there is always something to compare against.
Because every prediction records exactly what it was based on, old states can be rebuilt. A replay test rebuilt 32 of 32 saved states with matching hashes. It also caught a real bug where learned traits were leaking into states they shouldn't have been in.
A lab for comparing predictors
The /lab dashboard puts all six predictors side by side on accuracy, log loss, Brier score, calibration error, cost and latency. It also watches the rules above. If a prediction was ever shown before it was sealed, the dashboard would flag it.

Honest prompt optimization
I built an evaluation and prompt-optimization harness (GEPA) that runs from GitHub Actions, with spend caps, a check for answer leakage, and a held-out set. The first time it found an improvement in production, the gain didn't hold up on held-out data. So I changed the rule: the harness now reports an improvement only when it replicates.
Privacy that is tested
Tests try to infer a person's religion from answers they didn't consent to share, and they fail if any sensitive trait gets filled in that way. Another test checks that deleting a mimic really clears it from every store.
Where it stands
Mimic is deployed and invite-only. The first production report covers 4 people who opted in and 265 questions, which is too few to claim anything about accuracy yet, and the README says so. What's done is the part that makes a later claim believable: sealed predictions, a baseline on every question, replayable state and a lab that shows the numbers.