Best value
—
At any given level of coding ability, what is the cheapest model variant that reaches it?
—
—
—
—
highest pts/$ that clears each floor
Cost on a log scale. Bars are 95% confidence intervals — two variants whose bars overlap are not distinguishable. Greyed points are dominated: something else is both better and cheaper.
— click a row for provenance.
Each model family keeps only its two newest generations. Older ones still appear on the source leaderboards; they are listed here rather than dropped without trace.
| # | Variant | Effort | Provider | Index | DeepSWE | Terminal-Bench | SWE-Atlas | Cost/task | Time/task | Points/$ |
|---|
A benchmark is excluded when two or more models score 85% or above, or when it has not been updated in 30 days. Excluded benchmarks are listed, not hidden. If the filter would leave fewer than two benchmarks, nothing is excluded and this panel says so.
A variant's index and cost are assembled from the best-performing harness on each benchmark. One row can therefore mix Claude Code on one benchmark with Codex on another. The row describes an upper bound achievable across harnesses — not a single reproducible configuration you can buy. The harness that supplied each number is shown on every benchmark cell and in the expanded row.
The index is the unweighted mean of the raw percentages, so no score changes retroactively when a new model tops a leaderboard. Cost is averaged only over benchmarks that publish one; SWE-Atlas contributes score and no cost. Time and steps are shown but never enter the index or points per dollar. Scores are never imputed. Costs may be inherited, and are tagged est. when they are.
This is not a capability leaderboard — Artificial Analysis does that, and this index uses the same three benchmarks, so the two are directly comparable.
Refreshing re-scrapes the three leaderboards and requires the maintainer's key. Without it you are seeing the cached snapshot and its true age.