- Coding agents initially felt magical but now feel like dialup due to reliability issues and infrastructure strain.
- Token usage is exploding (50x increase on OpenRouter) because agentic workflows consume ~1000x more tokens than chats.
- Current LLM speed (30-60 tok/s) is frustrating, while faster models (2000 tok/s) could enable more unsupervised parallel approaches.
- The future of coding agents may involve parallel attempts and automated evaluation, enabled by higher tok/s.
- Infrastructure scaling faces challenges from slower semiconductor improvements, leading to potential pricing model innovations like off-peak plans.
- Developers should stay curious and adapt; experienced developers often dismiss coding agents but benefit most from them.