AI coding agent benchmark — capability versus cost per task
How well do frontier models actually finish real engineering tasks, and what does each attempt cost? The numbers below are the public DeepSWE leaderboard (v1.1) by Datacurve — 50 configurations across 18 models on 113 agentic tasks, pass@1 over repeated runs with 95% confidence intervals, generated 2026-07-25. All figures belong to Datacurve; we cite them, we don't own them.
Why cite this one: we ran our own agentic evaluations against a private ~1M-line polyglot repository (Java, Go, Python, Vue, React, Next.js) and our ordering agrees with theirs — so this is the closest public, reproducible reference to what we see on real production code. Prices per model live on our pricing pages.
Best configuration per model
Each model at the reasoning effort that scored highest — the short answer to "which one should I use", with what that run costs and how much pass rate you get per dollar.
| model | best effort | pass@1 | cost/task | pts per $ |
|---|---|---|---|---|
| claude-opus-5 | max | 73.7% | $11.84 | 6.2 |
| gpt-5-6-sol | max | 72.7% | $8.39 | 8.7 |
| claude-fable-5 | xhigh | 69.9% | $13.41 | 5.2 |
| gpt-5-6-terra | max | 69.6% | $4.95 | 14.1 |
| kimi-k3 | max | 68.5% | $4.65 | 14.7 |
| gpt-5-6-luna | max | 67.2% | $3.03 | 22.2 |
| gpt-5-5 | xhigh | 67.0% | $7.23 | 9.3 |
| claude-opus-4-8 | max | 59.0% | $13.22 | 4.5 |
| claude-sonnet-5 | max | 53.8% | $26.40 | 2.0 |
| grok-4-5 | high | 53.8% | $2.42 | 22.2 |
| muse-spark-1-1 | xhigh | 53.3% | $2.36 | 22.6 |
| gpt-5-4 | xhigh | 51.8% | $5.65 | 9.2 |
| gemini-3-6-flash | high | 48.6% | $3.53 | 13.8 |
| glm-5-2 | max | 43.8% | $3.92 | 11.2 |
| gemini-3-5-flash | medium | 37.4% | $7.34 | 5.1 |
| kimi-k2-7-code | — | 30.5% | $2.82 | 10.8 |
| claude-sonnet-4-6 | high | 29.9% | $5.52 | 5.4 |
| gemini-3-1-pro-preview | high | 11.8% | $9.48 | 1.2 |
Capability versus cost per task
Hover or tab a model to isolate its effort curve — connected dots are the same model at different reasoning efforts. Outlined dots sit on the Pareto frontier.
Full leaderboard — 50 configurations
| model | effort | pass@1 | 95% CI | pass@4 | cost/task | steps |
|---|---|---|---|---|---|---|
| claude-opus-5 | max | 73.7% | 70–78% | 88.5% | $11.84 | 90.5 |
| claude-opus-5 | xhigh | 73.2% | 70–76% | 85.8% | $9.07 | 80 |
| claude-opus-5 | high | 72.8% | 71–75% | 87.6% | $6.08 | 64 |
| gpt-5-6-sol | max | 72.7% | 70–76% | 85.8% | $8.39 | 53 |
| gpt-5-6-sol | xhigh | 70.7% | 70–72% | 85.8% | $4.70 | 39 |
| claude-fable-5 | xhigh | 69.9% | 67–73% | 88.5% | $13.41 | 61.5 |
| claude-fable-5 | max | 69.7% | 66–74% | 84.1% | $21.63 | 79 |
| gpt-5-6-terra | max | 69.6% | 67–72% | 88.5% | $4.95 | 71 |
| gpt-5-6-sol | high | 69.4% | 68–71% | 86.7% | $3.47 | 32 |
| claude-opus-5 | medium | 68.9% | 68–70% | 89.4% | $3.29 | 43 |
| claude-fable-5 | high | 68.6% | 67–70% | 86.7% | $9.18 | 49 |
| kimi-k3 | max | 68.5% | 64–73% | 89.4% | $4.65 | 88 |
| gpt-5-6-luna | max | 67.2% | 63–71% | 90.3% | $3.03 | 92.5 |
| gpt-5-5 | xhigh | 67.0% | 61–74% | 88.5% | $7.23 | 76.5 |
| claude-fable-5 | medium | 65.4% | 61–70% | 83.2% | $6.09 | 41 |
| gpt-5-5 | high | 64.4% | 61–68% | 90.3% | $5.10 | 60 |
| gpt-5-6-sol | medium | 61.1% | 59–63% | 80.5% | $1.86 | 26 |
| gpt-5-6-terra | xhigh | 60.2% | 58–62% | 80.5% | $2.13 | 39 |
| claude-fable-5 | low | 59.6% | 57–62% | 81.4% | $3.76 | 29 |
| claude-opus-4-8 | max | 59.0% | 57–61% | 79.3% | $13.22 | 116 |
| claude-opus-5 | low | 58.1% | 56–60% | 85.0% | $1.66 | 29 |
| gpt-5-6-luna | xhigh | 56.9% | 55–59% | 78.8% | $1.54 | 63 |
| claude-opus-4-8 | xhigh | 54.4% | 51–58% | 80.5% | $8.01 | 90 |
| gpt-5-5 | medium | 54.0% | 51–57% | 77.9% | $2.75 | 43 |
| claude-sonnet-5 | max | 53.8% | 50–58% | 78.8% | $26.40 | 260 |
| gpt-5-6-terra | high | 53.8% | 49–58% | 80.5% | $1.13 | 31 |
| grok-4-5 | high | 53.8% | 51–56% | 77.9% | $2.42 | 56 |
| muse-spark-1-1 | xhigh | 53.3% | 50–56% | 79.7% | $2.36 | 86 |
| claude-opus-4-8 | high | 51.8% | 47–56% | 77.9% | $4.28 | 67 |
| gpt-5-4 | xhigh | 51.8% | 50–53% | 77.9% | $5.65 | 63 |
| claude-sonnet-5 | xhigh | 49.7% | 46–53% | 75.2% | $11.89 | 174 |
| claude-opus-4-8 | medium | 48.7% | 46–51% | 76.1% | $3.44 | 60 |
| gemini-3-6-flash | high | 48.6% | 44–54% | 76.1% | $3.53 | 95 |
| claude-sonnet-5 | high | 48.2% | 44–53% | 79.7% | $7.43 | 138 |
| gpt-5-6-sol | low | 45.4% | 43–48% | 71.7% | $1.07 | 21 |
| gpt-5-6-luna | high | 44.3% | 41–47% | 75.2% | $0.78 | 44 |
| glm-5-2 | max | 43.8% | 42–46% | 77.0% | $3.92 | 123 |
| claude-opus-4-8 | low | 40.8% | 39–42% | 68.1% | $2.29 | 47 |
| claude-sonnet-5 | medium | 39.8% | 37–43% | 64.6% | $4.08 | 100.5 |
| gemini-3-5-flash | medium | 37.4% | 36–39% | 66.4% | $7.34 | 82 |
| glm-5-2 | high | 36.3% | 32–41% | 68.1% | $2.84 | 112 |
| gpt-5-6-terra | medium | 35.1% | 32–38% | 60.2% | $0.58 | 24 |
| kimi-k2-7-code | — | 30.5% | 30–31% | 61.1% | $2.82 | 139 |
| claude-sonnet-5 | low | 30.5% | 29–32% | 57.5% | $2.19 | 70 |
| claude-sonnet-4-6 | high | 29.9% | 26–34% | 56.6% | $5.52 | 124 |
| gpt-5-5 | low | 27.0% | 25–29% | 47.8% | $1.20 | 27 |
| gpt-5-6-terra | low | 24.1% | 23–25% | 44.3% | $0.43 | 20 |
| gemini-3-1-pro-preview | high | 11.8% | 9–14% | 28.3% | $9.48 | 77 |
| gpt-5-6-luna | medium | 11.3% | 10–12% | 27.4% | $0.22 | 22 |
| gpt-5-6-luna | low | 1.6% | 1–2% | 4.4% | $0.07 | 12 |
What the numbers mean
pass@1 is attempt pass rate over scored rollout attempts. pass@4 is tasks with at least one passing rollout divided by tasks attempted. Context-window failures and agent timeouts are scored failures; provider/verifier/network errors are excluded. Efficiency aggregates are over every scored attempt.
Every DeepSWE rollout across imported Pier jobs, grouped by configuration (harness + model + reasoning effort)
Reasoning effort is a per-run setting, not a different model — and it moves the numbers
as much as switching vendors does. The widest case here is
gpt-5-5: 27.0%
at low for $1.20
versus 67.0% at
xhigh for $7.23 —
40.1 points of pass rate for
6.0× the bill, same
model. The practical lesson: tune effort before you switch vendors.
Best value above a 60% bar is
gpt-5-6-sol medium at
61.1% for $1.86
(32.8 points per dollar),
while the top score costs $11.84.
DeepSWE leaderboard by Datacurve (original source) our LLM API pricing comparison token cost calculator
The weekly e/acc newsletter
One email a week: what accelerated.
Frontier releases, price drops, compute buildouts — the week's acceleration in five minutes, sourced and numeric. Free.