- AISI has tracked frontier AI cyber capabilities since 2023, with closed-weight models consistently outperforming open-weight models.
- Open-weight models offer benefits like private hosting, customization, and collaboration, but also pose risks due to lack of safety measures after release.
- The cyber capability gap between open and closed models narrowed from 6-10 months in 2025 to 4-7 months for GLM-5.2 (June 2026).
- On narrow cyber tasks, GLM-5.2 performed comparably to closed models from 4-5 months earlier, while DeepSeek V4-Pro lagged by 5 months.
- On longer-horizon cyber ranges, GLM-5.2 matched Opus 4.5 (released 7 months prior), but the gap was larger than in narrow tasks.
- Open-weight models are cheaper (e.g., DeepSeek V4-Pro at $0.28/task vs. $12.50 for Opus 4.5) but have limited safeguards that are easily removed.
- AISI's testing may underestimate open-weight models' capability, and findings are specific to cyber, not generalizable to other domains.
- The narrowing gap underscores the urgency for cyber defenders to prepare as frontier capabilities may become widely available without safeguards.