Hasty Briefsbeta

Bilingual

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

21 days ago
  • SWE-2 is an advanced coding model that achieves 50.0% on FrontierCode 1.1 Main, close to Fable 5.1 but 64% cheaper, pushing the Pareto frontier of capability and cost.
  • It uses a novel RL algorithm that trains all reasoning-effort levels in a single run, scaled to a multi-trillion-parameter regime for the first time.
  • SWE-2 beats previous models like SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and is within a few points of GPT-6 Astra at a quarter of the cost.
  • Improvements include Pareto-informed cost penalties, length-weighted reward baselines for stable training, enhanced RL rollout serving with NVFP4/FP8 kernels and quantization-aware training, and tripled RL environments with hardened verifiers.
  • Behaviorally, SWE-2 is more efficient: higher intelligence leads to focused exploration, quicker first edits (median 18 steps vs 48 for SWE-1.7), better test coverage, resourcefulness, and verification discipline.
  • Trustworthiness evaluations show a 98% pass rate in propaganda/censorship tests and no significant framing effects in context-dependent vulnerability tasks across different customer identities and languages.