IA / Agents~60 · IA en attente

Sonnet 5.5 vs Opus 5.5: the Terminal-Bench "win" is Max vs Xhigh. Here's the fair comparison.

r/ClaudeAIu/Intelligent-Lynx-95329 septembre 2026

Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.

Résumé

Sonnet 5.5 dropped yesterday and I keep seeing "Sonnet beats Opus" on the strength of Terminal-Bench 4.0 (70.6% vs 66.4%). That comparison has Sonnet at Max effort and Opus at Xhigh. At the same Xhigh setting, Sonnet scores 61.5%, so Opus actually leads by about 5 points. The full picture from Anthropic's launch table…

Afficher le post original
Sonnet 5.5 dropped yesterday and I keep seeing "Sonnet beats Opus" on the strength of Terminal-Bench 4.0 (70.6% vs 66.4%). That comparison has Sonnet at Max effort and Opus at Xhigh. At the same Xhigh setting, Sonnet scores 61.5%, so Opus actually leads by about 5 points. The full picture from Anthropic's launch table: Benchmark Sonnet 5.5 Opus 5.5 Gap GDPval-AA (knowledge work) 1844 1846 tied AA-Briefcase (knowledge work) 1811 1822 Opus +11 OSWorld 2.1 (computer use) 80.1% 81.8% Opus +1.7 CursorBench 4.0 55.5% 57.8% Opus +2.3 Chartography 61.6% 64.4% Opus +2.8 HLE (with tools) 64.5% 67.7% Opus +3.2 FrontierCode 1.1 46.2% 54.4% Opus +8.2 Terminal-Bench 4.0 (same effort) 61.5% 66.4% Opus +4.9 Here's my take: • Docs, spreadsheets, research write-ups, computer use: effectively a tie. Sonnet costs half as much ($2/$10 vs $4/$20 per million input/output tokens) and is about 30% faster (TechCrunch (https://techcrunch.com/2026/09/28/anthropic-releases-sonnet-5-5-which-it-calls-a-significantly-cheaper-faster-work-partner/)). Easy pick. • Hard agentic coding and messy debugging: Opus is still meaningfully ahead, around 5 to 8 points on the harder coding benchmarks. Anthropic itself says Opus 5.5 stays "clearly stronger" on open-ended work (The Decoder (https://the-decoder.com/anthropics-claude-sonnet-5-5-nearly-matches-opus-5-5-on-benchmarks-while-costing-up-to-30-percent-less-per-task/)). • Subagents and high-volume work: Sonnet 5.5 is a free upgrade over Sonnet 5 at the same price. The jump is enormous (Terminal-Bench went from 10.3% to 70.6%). So Sonnet isn't better than Opus. It's very nearly Opus on everyday work at half the price, which is arguably the bigger deal. Anyone run them side by side on real repos yet? Curious whether the FrontierCode gap shows up in practice? Sources: • Anthropic: Introducing Claude Sonnet 5.5 (https://www.anthropic.com/claude-sonnet-5-5) • Kingy AI: Sonnet 5.5 vs Opus 5.5 (https://kingy.ai/blog/claude-sonnet-5-5-vs-opus-5-5/) • llm-stats: Opus 5.5 vs Sonnet 5.5 (https://llm-stats.com/models/compare/claude-opus-5-5-vs-claude-sonnet-5-5) • TechCrunch (https://techcrunch.com/2026/09/28/anthropic-releases-sonnet-5-5-which-it-calls-a-significantly-cheaper-faster-work-partner/) • The Decoder (https://the-decoder.com/anthropics-claude-sonnet-5-5-nearly-matches-opus-5-5-on-benchmarks-while-costing-up-to-30-percent-less-per-task/)
Sonnet 5.5 vs Opus 5.5: the Terminal-Bench "win" is Max vs Xhigh. Here's the fair comparison. · Claude Project Scanner