- Databricks developed an internal coding benchmark using real engineering tasks from their codebase, evaluating AI coding agents on performance and cost.
- The benchmark identified three distinct capability tiers among models, indicating that high-cost models aren't always necessary for common tasks, with models like GLM 5.2 offering competitive quality at lower costs.
- Token costs often poorly predict overall task expenses due to variations in reasoning efficiency, with differences in harness performance highlighting the importance of context management.
- The team emphasized the need for custom benchmarks over public ones like SWE-Bench to accurately reflect internal needs and ensure optimizations don't hinder developers.