Sonnet 5 vs Sonnet 4.6 vs Opus 4.8

How Anthropic's new Sonnet 5 stacks up against the previous Sonnet 4.6 and the flagship Opus 4.8 across agentic coding, multidisciplinary reasoning, computer use, and knowledge work. Sonnet 5 delivers a big jump over Sonnet 4.6 and closes most of the gap to Opus 4.8 — while remaining the mid-tier model. Higher is better on every benchmark below.

Benchmark Comparison
Benchmark Sonnet 5 ⭐ Sonnet 4.6 Opus 4.8 (reference)
💻 Agentic coding
SWE-bench Pro63.2%58.1%69.2%
Terminal-Bench 2.180.4%67.0%82.7%
🧠 Multidisciplinary reasoning
Humanity's Last Exam (no tools)43.2%34.6%49.8%
Humanity's Last Exam (with tools)57.4%46.8%57.9%
🖥️ Computer use
OSWorld-Verified81.2%78.5%83.4%
📊 Knowledge work
GDPval-AA v2 (Elo score)161813951615

Figures are percentages except GDPval-AA v2, which is an Elo-style rating. Opus 4.8 is shown for reference as the flagship tier. Sonnet 5's column is highlighted as the featured model.

A short read of the numbers

The headline story is generational improvement in the mid tier. Across every benchmark, Sonnet 5 beats Sonnet 4.6 — often by wide margins. On Terminal-Bench 2.1 it jumps from 67.0% to 80.4% (+13.4 points), and on Humanity's Last Exam it gains roughly 9–11 points both with and without tools. The GDPval-AA v2 knowledge-work rating leaps from 1395 to 1618, a 223-point Elo swing that reflects a substantially more capable model on real-world tasks.

More striking is how close Sonnet 5 gets to the flagship. On the "with tools" Humanity's Last Exam (57.4% vs 57.9%) and on GDPval-AA v2 (1618 vs 1615) it is effectively level with Opus 4.8 — and on knowledge work it nudges just ahead. On computer use (81.2% vs 83.4%) and Terminal-Bench (80.4% vs 82.7%) it trails by only about two points.

Where Opus 4.8 still holds a clear edge is the hardest raw-reasoning and agentic-coding tests: SWE-bench Pro (69.2% vs 63.2%, a 6-point gap) and no-tools Humanity's Last Exam (49.8% vs 43.2%). The takeaway: Sonnet 5 offers most of Opus 4.8's capability at the mid tier, making it the sensible default for everyday coding and knowledge work — while Opus 4.8 remains the pick when you need the extra margin on the toughest problems.