Coding Agent Value Index

At any given level of coding ability, what is the cheapest model variant that reaches it?

Filters
Provider

Best value

Cheapest qualifying

Fastest qualifying

Highest index

Best value by floor

highest pts/$ that clears each floor

    Index vs cost per task

    Cost on a log scale. Bars are 95% confidence intervals — two variants whose bars overlap are not distinguishable. Greyed points are dominated: something else is both better and cheaper.

    All variants

    — click a row for provenance.

    # Variant Effort Provider Index DeepSWE Terminal-Bench SWE-Atlas Cost/task Time/task Points/$

    Benchmarks in this index

    A benchmark is excluded when two or more models score 85% or above, or when it has not been updated in 30 days. Excluded benchmarks are listed, not hidden. If the filter would leave fewer than two benchmarks, nothing is excluded and this panel says so.

    What this row is, and is not

    A variant's index and cost are assembled from the best-performing harness on each benchmark. One row can therefore mix Claude Code on one benchmark with Codex on another. The row describes an upper bound achievable across harnesses — not a single reproducible configuration you can buy. The harness that supplied each number is shown on every benchmark cell and in the expanded row.

    The index is the unweighted mean of the raw percentages, so no score changes retroactively when a new model tops a leaderboard. Cost is averaged only over benchmarks that publish one; SWE-Atlas contributes score and no cost. Time and steps are shown but never enter the index or points per dollar. Scores are never imputed. Costs may be inherited, and are tagged est. when they are.

    This is not a capability leaderboard — Artificial Analysis does that, and this index uses the same three benchmarks, so the two are directly comparable.

    Data

    Refreshing re-scrapes the three leaderboards and requires the maintainer's key. Without it you are seeing the cached snapshot and its true age.