Hasty Briefsbeta

Bilingual

Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

20 days ago
  • SWE-2 is a MoE model with 2.8 trillion total parameters and 104 billion active per token, built on Kimi K3 base with Cognition's post-training and reinforcement learning scaling into the multi-trillion-parameter regime.
  • Serving stack uses NVFP4 and FP8 kernels with quantization-aware training; FP8 handles K, Q, V, and score computations in MLA layers; SpecForge draft model yields 15% longer accept lengths; prefill delayer improves TPM per GPU and tokens/sec per request by 10-20% at the cost of higher TTFT.
  • Effort levels: medium (53 mean steps), high (80), max (98); medium achieves higher FrontierCode score than SWE-1.7 with 58% fewer turns and 81% lower cost, with first real edit at median step 18 vs 48 for SWE-1.7.
  • Benchmarks (Cognition self-reported): FrontierCode 50.0, DeepSWE 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3; FrontierCode trails Claude Fable 5.1 (50.9) and GPT-6 Astra (53.3) but at 64% lower cost than Fable and quarter of Astra's cost; Terminal-Bench 4.0 shows significant gap in long-horizon agentic work.
  • SWE-2 has proprietary weights, no local download; available in Devin Desktop and CLI, rolling out to Devin Web and Fusion; no per-token API; cost-per-task comparisons are vendor-reported and pending independent replication.