Choosing a model usually means four browser tabs: a leaderboard to rank it, the vendor's page to check the numbers, a pricing table, and a forum thread about whether it fits on your GPU. ModelBeats puts all four in one place, and every score links back to where it came from.
It's built for developers deciding between hosted APIs and open-weight models they could run themselves.

What you can do with it
You can sort the full ranking by overall rating or by any benchmark. You can filter the model explorer down to what runs on your GPU. The compare page puts up to four models side by side with their specs, every benchmark they share, a capability radar and a short verdict.

The hardware page answers "what can I run?" for a given GPU or Mac. The fine-tune page ranks open models by what you can train on your hardware, with an estimated cost for each method: QLoRA, LoRA or a full fine-tune, using SFT, DPO or RL.


Every filter, sort and comparison lives in the URL, so any view can be shared as a link.
How the ranking works
Vendors report different benchmarks, so there's no single test every model has taken. I treat every benchmark two models both report as a head-to-head match between them, then fit a Bradley-Terry model to all of those results. It's the model underneath Elo-style ratings, so I call the result Frontier Elo.
Some pairs of models share a lot of benchmarks in one category, like math. Without a correction, that category would decide the ranking. So each match is weighted by one over the square root of how many benchmarks the pair shares in that category. A model also needs at least 3 benchmarks across at least 2 categories before it gets a rank.
This wasn't my first design. I started with a min-max index, replaced it with Bradley-Terry, and then added the category weighting. Each change came with sensitivity tests: dropping one category at a time, dropping one benchmark at a time, and checking the result doesn't depend on the tuning constant. I wrote down the alternatives I rejected and why. The fitting code is unit-tested against an independent Newton-Raphson solver.

Where the data comes from
The repository is the database. Every result row must include a source URL, and the build fails without one. A script turns the research files into the dataset, validates it, computes the scores, and writes a content-hashed snapshot that gets bundled into the site at build time. The computed scores are committed to git, so a change in the rankings shows up in code review.
Collecting 5,969 results by hand wasn't realistic, so I ran the research with AI agents in batches and then checked their work. There were 107 research batches and a completeness sweep across the whole catalog. After that came two full re-verification passes against primary sources. In one of them, 838 results were checked, 35 were corrected and 9 were thrown out. The search covered Chinese and Japanese sources too. Every model that still has no results has a written reason.
Under the hood
The site is TanStack Start with React 19 running on a single Cloudflare Worker. It had a D1 database and KV at first. I removed both once it was clear the data could be bundled into the Worker at build time, which removed a whole class of moving parts.
CI runs type checks, linting, unit tests and a build, then enforces size budgets: the largest client chunk has to stay under 200 KB gzipped and the data snapshot under 1.5 MB. Playwright and axe accessibility tests run against a production-like preview, and deploys only go out when all of it passes.