Hasty Briefsbeta

Bilingual

Kimi K3: second only to Fable 5 on AA-Briefcase

7 hours ago
  • Kimi K3 ranks second only to Claude Fable 5 on the AA-Briefcase benchmark, with an Elo of 1543, a significant improvement over its predecessor Kimi K2.6.
  • It shows strong objective and analytical performance (rubric pass rate 51%, analytical Elo 1754) but comparatively weaker presentation quality (presentation Elo 1471).
  • The model is expensive to run, costing $10.57 per task on average, with high token usage and 83 turns per task, leading to an average time of 56.4 minutes per task.
  • Other recent frontier model launches include Grok 4.5, GPT-5.6, Muse Spark 1.1, and Kimi K3 itself, with six labs now having models above 50 on the AI Index.