{"contract":"guth-news-publication-v1","article":{"article_id":"7240b8fc-e615-4018-bb4c-73cfe594b542","revision":1,"slug":"sonnet-5-5-nears-opus-5-5-on-benchmarks-but-uses-more-tokens-7240b8fc","title":"Sonnet 5.5 nears Opus 5.5 on benchmarks but uses more tokens","summary":"Artificial Analysis says Sonnet 5.5 reaches near-Opus performance at maximum effort, with the highest output-token use it has measured.","body":"Artificial Analysis ranks Claude Sonnet 5.5 second on its Intelligence Index at maximum effort, with a score of 56, two points behind Opus 5.5 at maximum effort. The report says the model gained 18 points over Sonnet 5, while its listed input and output rates remain unchanged from Sonnet 5’s latest pricing. Across several agentic and knowledge-work evaluations, the evaluator found results close to Opus 5.5, but said Sonnet 5.5 used considerably more output tokens to reach that level.\n\nOn Terminal-Bench 4.0, Sonnet 5.5 scored 64%, compared with 60% for Opus 5.5 and GPT-6 Astra. The report also found near-parity with Opus 5.5 on AA-Briefcase, GDPval-AA and AutomationBench-AA, with scores of 1811 versus 1822 Elo, 1844 versus 1846 Elo, and 71% versus 70%, respectively. On Terminal-Bench-Science, which is not part of the Intelligence Index, Sonnet 5.5 scored 53%, behind GPT-6 Astra and Opus 5.5.\n\nThe report’s main qualification is resource use: at maximum effort, Sonnet 5.5 generated about 193,000 output tokens per Intelligence Index task, the highest amount Artificial Analysis says it has measured. That is about 60% above Opus 5.5 at maximum effort and Sonnet 5 at maximum effort, and roughly seven times GPT-6 Astra at maximum effort. The evaluator estimates a cost of $7.60 per task, about 50% more than Sonnet 5, despite matching Sonnet 5’s latest rates of $2 per million input tokens and $10 per million output tokens.\n\nArtificial Analysis says lower-effort settings offer different performance and token-use tradeoffs; it found some GPT-6 Astra or Sol configurations delivered equivalent performance at lower cost. The report describes maximum effort as the most competitive Sonnet 5.5 setting on its intelligence-versus-cost comparison, but still places it narrowly behind GPT-6 Sol. For builders, the findings suggest benchmark scores alone may not capture the cost of running a model: output volume and effort setting can also affect task economics.\n\nThe evaluations used a pre-release deployment with a bug that could affect requests using structured outputs. Anthropic fixed the issue for the public release, and the evaluator expects minimal change or slightly understated performance, while saying it will rerun relevant tests. Sonnet 5.5 retains a one-million-token context window with image and text input, and offers five effort levels: low, medium, high, xhigh and max.","content_kind":"author_paraphrase","explanation":{"feature":"Artificial Analysis says Sonnet 5.5 reaches near-Opus performance at maximum effort, with the highest output-token use it has measured.","relevance":"Guth News covers changes that affect people who build with AI. Read the cited primary sources for the full details.","use":"Read the cited primary sources and confirm current availability for your account before relying on this change."},"announcement_date":null,"published_at":"2026-09-29T12:04:29.932Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-09-29T12:04:29.355Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/7240b8fc-e615-4018-bb4c-73cfe594b542","claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"Artificial Analysis","url":"https://artificialanalysis.ai/articles/claude-sonnet-5-5","fetched_at":"2026-09-29T11:14:42.746Z","sha256":"dedefc2fb863ab524231bccbef2cce2a629bc17d7e9ca462d9b4816fce639123"}],"receipt":{"receipt_id":"230dd302-79c7-4efe-aac6-62144bbf455c","envelope_sha256":"637d9a76da0be28b8d225c97d52f5c1cd93cd69411aeeff35e2e53f71805eeea"},"canonical_url":"https://news.guthlabs.ai/articles/sonnet-5-5-nears-opus-5-5-on-benchmarks-but-uses-more-tokens-7240b8fc"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-09-29T12:04:29.932Z","reviewed_at":"2026-09-29T12:04:29.355Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Sonnet 5.5 nears Opus 5.5 on benchmarks but uses more tokens","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/sonnet-5-5-nears-opus-5-5-on-benchmarks-but-uses-more-tokens-7240b8fc?revision=1"}]}