e-acc.ai

AI coding agent benchmark — capability versus cost per task

How well do frontier models actually finish real engineering tasks, and what does each attempt cost? The numbers below are the public DeepSWE leaderboard (v1.1) by Datacurve — 50 configurations across 18 models on 113 agentic tasks, pass@1 over repeated runs with 95% confidence intervals, generated 2026-07-25. All figures belong to Datacurve; we cite them, we don't own them.

Why cite this one: we ran our own agentic evaluations against a private ~1M-line polyglot repository (Java, Go, Python, Vue, React, Next.js) and our ordering agrees with theirs — so this is the closest public, reproducible reference to what we see on real production code. Prices per model live on our pricing pages.

Best configuration per model

Each model at the reasoning effort that scored highest — the short answer to "which one should I use", with what that run costs and how much pass rate you get per dollar.

model best effort pass@1 cost/task pts per $
claude-opus-5 max 73.7% $11.84 6.2
gpt-5-6-sol max 72.7% $8.39 8.7
claude-fable-5 xhigh 69.9% $13.41 5.2
gpt-5-6-terra max 69.6% $4.95 14.1
kimi-k3 max 68.5% $4.65 14.7
gpt-5-6-luna max 67.2% $3.03 22.2
gpt-5-5 xhigh 67.0% $7.23 9.3
claude-opus-4-8 max 59.0% $13.22 4.5
claude-sonnet-5 max 53.8% $26.40 2.0
grok-4-5 high 53.8% $2.42 22.2
muse-spark-1-1 xhigh 53.3% $2.36 22.6
gpt-5-4 xhigh 51.8% $5.65 9.2
gemini-3-6-flash high 48.6% $3.53 13.8
glm-5-2 max 43.8% $3.92 11.2
gemini-3-5-flash medium 37.4% $7.34 5.1
kimi-k2-7-code 30.5% $2.82 10.8
claude-sonnet-4-6 high 29.9% $5.52 5.4
gemini-3-1-pro-preview high 11.8% $9.48 1.2

Capability versus cost per task

0% 20% 40% 60% 80% $0.5 $1 $2 $5 $10 $20 mean cost per task, log scale claude-opus-5 low — 58.1% at $1.66 claude-opus-5 medium — 68.9% at $3.29 claude-opus-5 high — 72.8% at $6.08 claude-opus-5 xhigh — 73.2% at $9.07 claude-opus-5 max — 73.7% at $11.84 claude-opus-5 gpt-5-6-sol low — 45.4% at $1.07 gpt-5-6-sol medium — 61.1% at $1.86 gpt-5-6-sol high — 69.4% at $3.47 gpt-5-6-sol xhigh — 70.7% at $4.70 gpt-5-6-sol max — 72.7% at $8.39 gpt-5-6-sol claude-fable-5 low — 59.6% at $3.76 claude-fable-5 medium — 65.4% at $6.09 claude-fable-5 high — 68.6% at $9.18 claude-fable-5 xhigh — 69.9% at $13.41 claude-fable-5 max — 69.7% at $21.63 gpt-5-6-terra low — 24.1% at $0.43 gpt-5-6-terra medium — 35.1% at $0.58 gpt-5-6-terra high — 53.8% at $1.13 gpt-5-6-terra xhigh — 60.2% at $2.13 gpt-5-6-terra max — 69.6% at $4.95 kimi-k3 max — 68.5% at $4.65 kimi-k3 gpt-5-6-luna low — 1.6% at $0.07 gpt-5-6-luna medium — 11.3% at $0.22 gpt-5-6-luna high — 44.3% at $0.78 gpt-5-6-luna xhigh — 56.9% at $1.54 gpt-5-6-luna max — 67.2% at $3.03 gpt-5-5 low — 27.0% at $1.20 gpt-5-5 medium — 54.0% at $2.75 gpt-5-5 high — 64.4% at $5.10 gpt-5-5 xhigh — 67.0% at $7.23 claude-opus-4-8 low — 40.8% at $2.29 claude-opus-4-8 medium — 48.7% at $3.44 claude-opus-4-8 high — 51.8% at $4.28 claude-opus-4-8 xhigh — 54.4% at $8.01 claude-opus-4-8 max — 59.0% at $13.22 claude-sonnet-5 low — 30.5% at $2.19 claude-sonnet-5 medium — 39.8% at $4.08 claude-sonnet-5 high — 48.2% at $7.43 claude-sonnet-5 xhigh — 49.7% at $11.89 claude-sonnet-5 max — 53.8% at $26.40 grok-4-5 high — 53.8% at $2.42 grok-4-5 muse-spark-1-1 xhigh — 53.3% at $2.36 muse-spark-1-1 gpt-5-4 xhigh — 51.8% at $5.65 gemini-3-6-flash high — 48.6% at $3.53 gemini-3-6-flash glm-5-2 high — 36.3% at $2.84 glm-5-2 max — 43.8% at $3.92 glm-5-2 gemini-3-5-flash medium — 37.4% at $7.34 kimi-k2-7-code — 30.5% at $2.82 claude-sonnet-4-6 high — 29.9% at $5.52 gemini-3-1-pro-preview high — 11.8% at $9.48

Hover or tab a model to isolate its effort curve — connected dots are the same model at different reasoning efforts. Outlined dots sit on the Pareto frontier.

Source: DeepSWE v1.1 (Datacurve), generated 2026-07-25.

Full leaderboard — 50 configurations

model effort pass@1 95% CI pass@4 cost/task steps
claude-opus-5 max 73.7% 70–78% 88.5% $11.84 90.5
claude-opus-5 xhigh 73.2% 70–76% 85.8% $9.07 80
claude-opus-5 high 72.8% 71–75% 87.6% $6.08 64
gpt-5-6-sol max 72.7% 70–76% 85.8% $8.39 53
gpt-5-6-sol xhigh 70.7% 70–72% 85.8% $4.70 39
claude-fable-5 xhigh 69.9% 67–73% 88.5% $13.41 61.5
claude-fable-5 max 69.7% 66–74% 84.1% $21.63 79
gpt-5-6-terra max 69.6% 67–72% 88.5% $4.95 71
gpt-5-6-sol high 69.4% 68–71% 86.7% $3.47 32
claude-opus-5 medium 68.9% 68–70% 89.4% $3.29 43
claude-fable-5 high 68.6% 67–70% 86.7% $9.18 49
kimi-k3 max 68.5% 64–73% 89.4% $4.65 88
gpt-5-6-luna max 67.2% 63–71% 90.3% $3.03 92.5
gpt-5-5 xhigh 67.0% 61–74% 88.5% $7.23 76.5
claude-fable-5 medium 65.4% 61–70% 83.2% $6.09 41
gpt-5-5 high 64.4% 61–68% 90.3% $5.10 60
gpt-5-6-sol medium 61.1% 59–63% 80.5% $1.86 26
gpt-5-6-terra xhigh 60.2% 58–62% 80.5% $2.13 39
claude-fable-5 low 59.6% 57–62% 81.4% $3.76 29
claude-opus-4-8 max 59.0% 57–61% 79.3% $13.22 116
claude-opus-5 low 58.1% 56–60% 85.0% $1.66 29
gpt-5-6-luna xhigh 56.9% 55–59% 78.8% $1.54 63
claude-opus-4-8 xhigh 54.4% 51–58% 80.5% $8.01 90
gpt-5-5 medium 54.0% 51–57% 77.9% $2.75 43
claude-sonnet-5 max 53.8% 50–58% 78.8% $26.40 260
gpt-5-6-terra high 53.8% 49–58% 80.5% $1.13 31
grok-4-5 high 53.8% 51–56% 77.9% $2.42 56
muse-spark-1-1 xhigh 53.3% 50–56% 79.7% $2.36 86
claude-opus-4-8 high 51.8% 47–56% 77.9% $4.28 67
gpt-5-4 xhigh 51.8% 50–53% 77.9% $5.65 63
claude-sonnet-5 xhigh 49.7% 46–53% 75.2% $11.89 174
claude-opus-4-8 medium 48.7% 46–51% 76.1% $3.44 60
gemini-3-6-flash high 48.6% 44–54% 76.1% $3.53 95
claude-sonnet-5 high 48.2% 44–53% 79.7% $7.43 138
gpt-5-6-sol low 45.4% 43–48% 71.7% $1.07 21
gpt-5-6-luna high 44.3% 41–47% 75.2% $0.78 44
glm-5-2 max 43.8% 42–46% 77.0% $3.92 123
claude-opus-4-8 low 40.8% 39–42% 68.1% $2.29 47
claude-sonnet-5 medium 39.8% 37–43% 64.6% $4.08 100.5
gemini-3-5-flash medium 37.4% 36–39% 66.4% $7.34 82
glm-5-2 high 36.3% 32–41% 68.1% $2.84 112
gpt-5-6-terra medium 35.1% 32–38% 60.2% $0.58 24
kimi-k2-7-code 30.5% 30–31% 61.1% $2.82 139
claude-sonnet-5 low 30.5% 29–32% 57.5% $2.19 70
claude-sonnet-4-6 high 29.9% 26–34% 56.6% $5.52 124
gpt-5-5 low 27.0% 25–29% 47.8% $1.20 27
gpt-5-6-terra low 24.1% 23–25% 44.3% $0.43 20
gemini-3-1-pro-preview high 11.8% 9–14% 28.3% $9.48 77
gpt-5-6-luna medium 11.3% 10–12% 27.4% $0.22 22
gpt-5-6-luna low 1.6% 1–2% 4.4% $0.07 12

What the numbers mean

pass@1 is attempt pass rate over scored rollout attempts. pass@4 is tasks with at least one passing rollout divided by tasks attempted. Context-window failures and agent timeouts are scored failures; provider/verifier/network errors are excluded. Efficiency aggregates are over every scored attempt.

Every DeepSWE rollout across imported Pier jobs, grouped by configuration (harness + model + reasoning effort)

Reasoning effort is a per-run setting, not a different model — and it moves the numbers as much as switching vendors does. The widest case here is gpt-5-5: 27.0% at low for $1.20 versus 67.0% at xhigh for $7.23 — 40.1 points of pass rate for 6.0× the bill, same model. The practical lesson: tune effort before you switch vendors. Best value above a 60% bar is gpt-5-6-sol medium at 61.1% for $1.86 (32.8 points per dollar), while the top score costs $11.84.

The weekly e/acc newsletter

One email a week: what accelerated.

Frontier releases, price drops, compute buildouts — the week's acceleration in five minutes, sourced and numeric. Free.