Hasty Briefsbeta

Bilingual

How Far Behind the Frontier Are Leading Open Weight Models on Cyber?

8 hours ago
  • AISI has tracked frontier AI cyber capabilities since 2023, with closed-weight models consistently outperforming open-weight models.
  • Open-weight models offer benefits like private hosting, customization, and collaboration, but also pose risks due to lack of safety measures after release.
  • The cyber capability gap between open and closed models narrowed from 6-10 months in 2025 to 4-7 months for GLM-5.2 (June 2026).
  • On narrow cyber tasks, GLM-5.2 performed comparably to closed models from 4-5 months earlier, while DeepSeek V4-Pro lagged by 5 months.
  • On longer-horizon cyber ranges, GLM-5.2 matched Opus 4.5 (released 7 months prior), but the gap was larger than in narrow tasks.
  • Open-weight models are cheaper (e.g., DeepSeek V4-Pro at $0.28/task vs. $12.50 for Opus 4.5) but have limited safeguards that are easily removed.
  • AISI's testing may underestimate open-weight models' capability, and findings are specific to cyber, not generalizable to other domains.
  • The narrowing gap underscores the urgency for cyber defenders to prepare as frontier capabilities may become widely available without safeguards.