SWE-bench Verified
Resolve real GitHub issues in an actual repo — the reference test for agentic coding.
FrontierBenchmarks measure AI models' abilities — code, reasoning, knowledge. They're useful but tricky. Here's what they really measure, how to read them, and where to track Claude's scores as releases land (no invented numbers: exact values change with every version).
Source: llm-stats.com · as of August 3, 2026
Note: this leaderboard does not yet include Claude Opus 5 (released 24 July 2026) — its source has not scored it. Anthropic reports roughly 96% on SWE-bench Verified, a figure from a different methodology, so it is not directly comparable to the rows above.
| Benchmark | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|
| SWE-bench ProAgentic coding, uncontaminated | 69.2% | 58.6% | 54.2% |
| Terminal-Bench 2.1Autonomy in a terminal | 74.6% | 78.2% | 70.3% |
| OSWorld-VerifiedOperating a computer | 83.4% | 78.7% | 76.2% |
| Humanity's Last ExamExpert questions (with tools) | 57.9% | 52.2% | 51.4% |
| Finance Agent v2Financial agent | 53.9% | 51.8% | 43.0% |
Source: Model card Anthropic — Claude Opus 4.8 · vendor conditions, as of August 3, 2026
These numbers are real and dated, not fixed: a score depends on conditions (with/without tools, harness, version) and moves with every release. Always compare at equal source and date — and test on your own task.
Resolve real GitHub issues in an actual repo — the reference test for agentic coding.
FrontierA hardened SWE-bench on actively maintained repos, with no public solution leakage.
FrontierComplete real terminal tasks (install, configure, debug) end to end.
FrontierRecent coding problems published after training — built to resist contamination.
FrontierPhD-level science questions, “Google-proof” — impossible to just look up.
FrontierThousands of expert, multi-domain questions written to stay hard for a long time.
FrontierA harder MMLU: more options and reworked questions to restore a gap between models.
ActiveUS math-olympiad problems — high-level mathematical reasoning.
FrontierCollege-level reasoning over image + text (figures, charts, diagrams).
ActiveDrive a real computer (mouse, keyboard, apps) to complete tasks — computer use.
FrontierWrite a correct function from a docstring. Historic, now near-ceiling.
SaturatedMultiple-choice across 57 fields (law, medicine, history…). The knowledge staple, now saturated.
SaturatedA few references recur: SWE-bench (solving real software bugs, key for agentic coding), MMLU and MMLU-Pro (general knowledge), GPQA (expert-level scientific reasoning), MATH and GSM8K (math), HumanEval (code generation). Each lights up a different facet — none alone captures 'intelligence'.
An isolated score often lies. Beware data contamination (the test may have leaked into training), conditions (with or without tools, which prompt), and which versions are compared. A model can top a benchmark and disappoint on your task. The best test is your own.
The new generation measures tool use and autonomy: SWE-bench Verified, TAU-bench, agent benchmarks. This is where the future is decided, and where models built for action — like those behind Claude Code — are judged.
Numbers shift with every model release. Rather than freezing a quickly outdated ranking, follow official announcements and our feed, Models category, which relays results as they come.
SWE-bench (and its Verified variant) is the reference for agentic coding, since it measures solving real software problems, not just isolated snippets.
It depends on the benchmark and versions compared: Claude is regularly on top for agentic coding and reasoning, but no model dominates everywhere. Check up-to-date scores.
In Anthropic's official announcements and in our feed (Models category), which relays benchmarks at each release.
With caution: data contamination, test conditions and version choices can skew the reading. A score is a hint, not a truth.
Claude News is published by Héra SASU. Independent media, not affiliated with Anthropic.