Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Index
12 hours ago
- Anthropic launched Claude Sonnet 5.5, scoring 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 but with the highest output tokens per task ever measured.
- At maximum effort, Sonnet 5.5 gains 18 points over Sonnet 5 and matches Opus 5.5 on agentic benchmarks like Terminal-Bench 4.0 and AA-Briefcase, though using significantly more tokens.
- The model is priced identically to Sonnet 5 at $0.2/$2/$10 per million cache/input/output tokens, but its cost per task is about 50% higher due to heavy token usage (~193k output tokens per task).
- Sonnet 5.5 lags behind Opus 5.5 on factual knowledge (54% vs 66%) and scientific reasoning (6 points lower on Humanity's Last Exam and SciCode), but has a lower hallucination rate.
- Anthropic fixed a pre-release bug affecting structured outputs; public release performance is expected to be slightly improved.
- The model maintains a 1 million token context window and offers five effort settings (low to max); at lower efforts, GPT-6 Sol provides better performance per token.