Hasty Briefsbeta

Bilingual

Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Index

12 hours ago
  • Anthropic launched Claude Sonnet 5.5, scoring 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 but with the highest output tokens per task ever measured.
  • At maximum effort, Sonnet 5.5 gains 18 points over Sonnet 5 and matches Opus 5.5 on agentic benchmarks like Terminal-Bench 4.0 and AA-Briefcase, though using significantly more tokens.
  • The model is priced identically to Sonnet 5 at $0.2/$2/$10 per million cache/input/output tokens, but its cost per task is about 50% higher due to heavy token usage (~193k output tokens per task).
  • Sonnet 5.5 lags behind Opus 5.5 on factual knowledge (54% vs 66%) and scientific reasoning (6 points lower on Humanity's Last Exam and SciCode), but has a lower hallucination rate.
  • Anthropic fixed a pre-release bug affecting structured outputs; public release performance is expected to be slightly improved.
  • The model maintains a 1 million token context window and offers five effort settings (low to max); at lower efforts, GPT-6 Sol provides better performance per token.