Hasty Briefsbeta

Bilingual

UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

11 hours ago
  • UK AISI and US CAISI jointly evaluated Moonshot AI's Kimi K3 model, focusing on its cyber capabilities.
  • Kimi K3 outperforms GLM-5.2 (32% vs 24%) on exploit development benchmark ExploitBench, but lags behind top US models.
  • Kimi K3 failed to achieve arbitrary code execution (ACE) on any of 41 ExploitBench tasks, while top US models achieved ACE on 20/41.
  • On the TLO cyber range (32-step simulated attack), Kimi K3 reached step 17 on average, compared to 28.5 steps for top US models.
  • Kimi K3 successfully completed TLO in 1 out of 10 attempts within a 100M token limit, indicating capability to attack weakly defended systems.
  • Kimi K3's overall cyber capability confidence interval is larger due to estimation from a single benchmark (ExploitBench).