Sonnet 5 vs Sonnet 4.6 vs Opus 4.8
How Anthropic's new Sonnet 5 stacks up against the previous Sonnet 4.6 and the flagship Opus 4.8 across agentic coding, multidisciplinary reasoning, computer use, and knowledge work. Sonnet 5 delivers a big jump over Sonnet 4.6 and closes most of the gap to Opus 4.8 — while remaining the mid-tier model. Higher is better on every benchmark below.
Benchmark Comparison
| Benchmark | Sonnet 5 ⭐ | Sonnet 4.6 | Opus 4.8 (reference) |
|---|---|---|---|
| 💻 Agentic coding | |||
| SWE-bench Pro | 63.2% | 58.1% | 69.2% |
| Terminal-Bench 2.1 | 80.4% | 67.0% | 82.7% |
| 🧠 Multidisciplinary reasoning | |||
| Humanity's Last Exam (no tools) | 43.2% | 34.6% | 49.8% |
| Humanity's Last Exam (with tools) | 57.4% | 46.8% | 57.9% |
| 🖥️ Computer use | |||
| OSWorld-Verified | 81.2% | 78.5% | 83.4% |
| 📊 Knowledge work | |||
| GDPval-AA v2 (Elo score) | 1618 | 1395 | 1615 |
Figures are percentages except GDPval-AA v2, which is an Elo-style rating. Opus 4.8 is shown for reference as the flagship tier. Sonnet 5's column is highlighted as the featured model.
A short read of the numbers
The headline story is generational improvement in the mid tier. Across every benchmark, Sonnet 5 beats Sonnet 4.6 — often by wide margins. On Terminal-Bench 2.1 it jumps from 67.0% to 80.4% (+13.4 points), and on Humanity's Last Exam it gains roughly 9–11 points both with and without tools. The GDPval-AA v2 knowledge-work rating leaps from 1395 to 1618, a 223-point Elo swing that reflects a substantially more capable model on real-world tasks.
More striking is how close Sonnet 5 gets to the flagship. On the "with tools" Humanity's Last Exam (57.4% vs 57.9%) and on GDPval-AA v2 (1618 vs 1615) it is effectively level with Opus 4.8 — and on knowledge work it nudges just ahead. On computer use (81.2% vs 83.4%) and Terminal-Bench (80.4% vs 82.7%) it trails by only about two points.
Where Opus 4.8 still holds a clear edge is the hardest raw-reasoning and agentic-coding tests: SWE-bench Pro (69.2% vs 63.2%, a 6-point gap) and no-tools Humanity's Last Exam (49.8% vs 43.2%). The takeaway: Sonnet 5 offers most of Opus 4.8's capability at the mid tier, making it the sensible default for everyday coding and knowledge work — while Opus 4.8 remains the pick when you need the extra margin on the toughest problems.