Kimi K3: second only to Fable 5 on AA-Briefcase
7 hours ago
- Kimi K3 ranks second only to Claude Fable 5 on the AA-Briefcase benchmark, with an Elo of 1543, a significant improvement over its predecessor Kimi K2.6.
- It shows strong objective and analytical performance (rubric pass rate 51%, analytical Elo 1754) but comparatively weaker presentation quality (presentation Elo 1471).
- The model is expensive to run, costing $10.57 per task on average, with high token usage and 83 turns per task, leading to an average time of 56.4 minutes per task.
- Other recent frontier model launches include Grok 4.5, GPT-5.6, Muse Spark 1.1, and Kimi K3 itself, with six labs now having models above 50 on the AI Index.