Skip to content

How to Read 2026 Coding Benchmarks: SWE-bench, Terminal-Bench, DeepSWE, Aider Explained

Aug 26, 2026 1 min
TL;DR The same model can score 20 points apart on different harnesses, 32% of SWE-bench Pro verifier judgments were found to be wrong, and DeepSWE's 113 tasks make most models score zero. This guide decodes six major coding benchmarks — what they test, which are easy to game, and which ones you should care about.
Table of Contents
  1. SWE-bench: The Most Cited Coding Benchmark
    1. Three Versions, Very Different
    2. What to Watch For
  2. Terminal-Bench 2.1: Testing Agents in the Terminal
    1. Harness-Dependent Score Differences
  3. DeepSWE: The Hardest Benchmark — Most Models Score Zero
  4. Aider Polyglot: The Most Practical Benchmark
    1. Why Aider Is More "Real-World" Than SWE-bench
  5. LiveCodeBench: Competitive Programming ≠ Software Engineering
  6. HumanEval / MBPP: Legacy Benchmarks Past Their Expiry Date
  7. Spotting Benchmark Gaming
  8. Which Benchmark Should You Care About?
  9. References

🌏 中文版

Every new model launch comes with a wall of numbers: SWE-bench 86%, Terminal-Bench 68.5%, DeepSWE 56. But what do these numbers actually mean? Why does the same model get different scores in different reports? Which benchmarks are easy to game? This guide is for anyone who wants to read model announcements without wading through papers.

SWE-bench: The Most Cited Coding Benchmark

SWE-bench is the most widely cited coding benchmark. It tests models on real GitHub issues: given an issue description and a code repository, the model must write a patch that passes the tests.

Three Versions, Very Different

Original SWE-bench (2023): 2,294 tasks from 12 Python repos. Many tasks had vague issue descriptions or incomplete test cases.

SWE-bench Verified (2024): OpenAI funded 93 professional engineers to manually review and curate 500 high-quality tasks. Per CodingFleet's analysis, OpenAI themselves publicly abandoned Verified in February 2026 — 500 tasks was too few for statistical power, and it only covered Python. Yet it remains the most cited version because it has the most historical scores.

SWE-bench Pro (2025): Built by Scale AI — 1,865 tasks across 41 repos and 123 languages. A ground-up redesign that reflects polyglot reality. Per the same analysis, a DeepSWE audit found 32% of Pro's verifier judgments were wrong — a number that illustrates how hard it is to build a correct benchmark.

What to Watch For

  • Verified is nearly saturated: top models cluster at 80–86%, differences fall within confidence intervals
  • Pro is the future: multilingual, larger task pool, but less historical data for comparison
  • Harness matters: the model only provides reasoning; an agent harness (e.g., mini-SWE-agent) handles file operations and command execution. Different harnesses produce different scores

Terminal-Bench 2.1: Testing Agents in the Terminal

Terminal-Bench tests a model's ability to operate as a terminal agent — not just writing code, but system administration, data processing, and environment configuration.

Per the Terminal-Bench 2.1 announcement, version 2.1 fixed 28 of 89 tasks from 2.0 (external dependency drift, insufficient resource budgets, misspecified instructions). After the fix, Claude Code + Opus 4.6 jumped 12.1%.

Harness-Dependent Score Differences

This is the easiest source of misreadings. Take Ornith 1.5-35B as an example:

HarnessTerminal-Bench 2.1 Score
Terminus-267.8
Claude Code harness68.5

A 0.7 gap is small, but some models differ by 10–20 points across harnesses. When reading Terminal-Bench scores, always check which harness was used. Artificial Analysis standardizes on Terminus 2 with 3 runs averaged.

DeepSWE: The Hardest Benchmark — Most Models Score Zero

DeepSWE, released by Datacurve in May 2026, is a "long-horizon software engineering" benchmark. 113 tasks across 91 repos and 5 languages, every task written from scratch rather than adapted from existing commits.

Key differences from SWE-bench:

  • Contamination-free: all tasks are original — no model could have seen solutions during pretraining
  • Genuinely hard: prompts are ~half the length of SWE-bench Pro's, yet solutions require 5.5× more code and ~2× more output tokens
  • Behavioral verification: tests check software behavior rather than implementation details

Per the official DeepSWE leaderboard (2026-08-20), the top score is Claude Opus 5 at 74%. MiniMax M2.5 scores 22%, Ornith 1.5-35B also 22% — while same-class models Qwen3.6-35B and Gemma 4-31B score zero. DeepSWE's discrimination power far exceeds saturated benchmarks.

Aider Polyglot: The Most Practical Benchmark

Aider's Polyglot benchmark uses 225 Exercism exercises across C++, Go, Java, JavaScript, Python, and Rust, testing coding ability in a pair-programming context.

Aider's unique feature is cost tracking. Per the official leaderboard, GPT-5 (high) scores 88% but costs $29.08 per run, while DeepSeek V3.2 scores 70.2% at just $0.88 — a 33× cost difference for 18 points. This enables cost-effectiveness analysis rather than pure score chasing.

Why Aider Is More "Real-World" Than SWE-bench

SWE-bench tests "here's a bug report, fix it." Aider tests "given existing code, follow instructions to write new features or refactor" — closer to how developers actually work with AI. But Aider's tasks are exercise-level, not real-repo complexity.

LiveCodeBench: Competitive Programming ≠ Software Engineering

LiveCodeBench uses LeetCode/Codeforces-level competitive programming problems. Nous Research's NousCoder-14B scored 67.87% Pass@1 on LiveCodeBench v6.

Key caveat: competitive programming and software engineering are different skills. Competitive programming tests algorithm design and edge-case handling; software engineering tests understanding large codebases, cross-file modifications, and test framework interaction. A model strong on LiveCodeBench but weak on SWE-bench is entirely possible.

HumanEval / MBPP: Legacy Benchmarks Past Their Expiry Date

HumanEval (164 tasks) and MBPP (974 tasks) were the first coding benchmarks. Top models exceed 95% on HumanEval — effectively saturated. They're still cited for one reason only: the longest historical record, useful for cross-era comparison.

In 2026, don't use HumanEval to judge model quality — it's like using elementary school math to evaluate PhD candidates.

Spotting Benchmark Gaming

TechniqueHow to Detect
Self-reported scoresNo independent third-party run. Check for verification on Artificial Analysis or SWE-bench official
Cherry-picking harnessSame model on different harnesses, only reporting the highest score. Proper practice: standardize harness or report all
Version cherry-pickingReporting SWE-bench Verified but not Pro, or Terminal-Bench 2.0 but not 2.1. Usually because the newer version scores lower
Training set contaminationModel saw benchmark solutions during pretraining. DeepSWE was created specifically to counter this
Best-of-NRunning many times and reporting the best instead of Pass@1. Proper benchmarks report Pass@1 with the number of runs averaged

Which Benchmark Should You Care About?

Your NeedLook AtWhy
Evaluating bug-fixing agentsSWE-bench ProLargest multilingual real-issue test set
Evaluating terminal agent capabilityTerminal-Bench 2.1Only benchmark specifically for terminal agents
Evaluating on the hardest tasksDeepSWEContamination-free, long-horizon, highest discrimination
Evaluating pair-programming cost-effectivenessAider PolyglotOnly benchmark that tracks cost alongside scores
Evaluating algorithmic abilityLiveCodeBenchMost current competitive programming benchmark
Cross-era comparison (2023–2026)HumanEvalLongest historical record, but saturated

One final piece of advice: never judge a model by a single benchmark. A model with high SWE-bench scores but zero on DeepSWE likely saw SWE-bench solutions during training. Cross-referencing multiple benchmarks is the correct approach.

References