13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
4 hours ago
- Fable 5 [high] model achieves the highest resolved rate at 64.5%, with a pass@5 of 78.4% and a cost of $4.40 per problem.
- GLM-5.2 [high] model has the highest pass@5 at 81.1%, while also maintaining a strong resolved rate of 62.9%.
- GPT-5.6 Sol [medium] offers a balance of high performance (62.3% resolved, 79.3% pass@5) with relatively low cost ($0.85) and token usage (605,340).
- MiMo V2.5 Pro model is the most cost-effective at $0.10 per problem, but has a lower resolved rate of 46.5%.
- MiniMax M3 model uses the most tokens per problem at 13,869,459, yet resolves only 47.2% of problems.
- Many models listed (from #18 onward) lack performance data, showing N/A for all metrics.