← All projects

OlympicDucks

An online arena for small simulated walking robots. Upload a control policy, or build a duck in the browser, and watch it climb a ranked ladder in sumo, soccer and a 25-meter sprint.

August 2026 - September 2026

sports: sumo, soccer, sprint
3
sports: sumo, soccer, sprint
house policies, 9 trained by me
20
house policies, 9 trained by me
tests across Python, TypeScript and the browser
432
tests across Python, TypeScript and the browser
dated decision records
182
dated decision records
Robotics
Reinforcement learning
MuJoCo
Python
Next.js
three.js
PostgreSQL
Fly.io

Pollen Robotics released the MicroDuck, a small open-source walking robot, in late August 2026. It shipped with walking policies trained by reinforcement learning, but there was nowhere to test one policy against another. I built OlympicDucks so people training these robots would have a public leaderboard they could trust.

Trust was the main design constraint. The browser never decides a result. Every ranked fight runs on my servers, and any result can be re-derived from what's stored.

The OlympicDucks home page with a live 3D replay of a soccer match.
The home page shows the best recent fight in each sport as a 3D replay you can scrub through.

How it works

You sign in and enter a duck in one of two ways. You can upload a neural network policy as an ONNX file, from the website or a small Python command-line tool. Or you can pick a stock walking style, tune a few strategy sliders, and watch a preview fight run right in your browser before you commit.

Once the server has checked your policy, the duck starts its climb. It fights the weakest house duck first and moves up one rung per win. The climb stops after two losses, after 12 fights, or when it reaches the top. After that the scheduler keeps finding it new opponents, starting with the ducks whose rating is least certain.

Each match gets a page with a 3D replay, a timeline of what happened, the rating change, and a shareable video.

The "How to reach #1" page explaining how ducks are ranked.
Ducks are ranked on a confirmed rating: the rating minus twice the remaining uncertainty.

Ratings use Glicko-2, and the ladder ranks ducks on a confirmed rating: their rating minus twice the uncertainty. A new duck on a hot streak doesn't jump straight to the top. It gets there once the arena is sure.

Under the hood

There are three parts, and they only talk through Postgres and file storage. There is no API between them.

The website is Next.js on Cloudflare Workers. It writes duck and job records and never runs a simulation. The browser preview runs MuJoCo compiled to WebAssembly in a sandboxed frame, and it's always labeled as a preview.

The workers are Python. Each one claims jobs from a queue built on Postgres itself (FOR UPDATE SKIP LOCKED with leases, heartbeats and retries), simulates the fight in MuJoCo, and stores a compressed trace of about 150 KB per minute of play. Each sport is a plugin, and its referee is a pure function of the recorded frames, so a stored replay can be judged again later.

The fleet runs on Fly.io. Fight machines and video-rendering machines are separate pools, and each scales on its own from the length of its queue. I looked at Temporal and Hatchet for the job system and decided plain Postgres was enough.

The hard parts

The same fight, bit for bit, on any CPU

This was the hardest problem in the project. The ONNX runtime picks different math kernels depending on the CPU, so the same walking policy made slightly different moves on different server CPUs. With a balancing robot, small differences grow fast, and the fights diverged.

I fixed it by writing my own policy evaluator in integer fixed-point math, so the result no longer depends on the processor. The ONNX runtime is still there, but only to check my evaluator against.

Seasons pinned to the physics

Every worker replays a set of known "golden" fights when it boots. If its results don't match exactly, it isn't allowed to rank anything. Each season records a fingerprint built from those golden results, so the fingerprint changes only when the physics really changes. A separate test re-simulates a finished match and compares the stored trace hash byte for byte.

Two renderers that match

The browser draws replays with three.js. The server makes share images and videos with MuJoCo's own renderer. I tuned the two by hand so a fight looks the same in the video as it does on the site.

Security tested by attacking it

I ran four adversarial security reviews. Each fix is shown to work by running the attack against it, then removing the fix and checking the attack succeeds again.

Credits

The duck meshes are derived from Pollen Robotics' Microduck and carry a CC BY-NC-SA 4.0 license, so OlympicDucks is non-commercial. Of the 20 house policies, 11 come from Pollen's public release and I trained the other 9. The code is private.