One month coding with GLM 5.3 Flash
5 hours ago
- The challenge to use only GLM 5.3 Flash for September succeeded only in the first half; the second half saw 1B tokens used on other models.
- Vibe coding an MCP server prototype with the 'wrong' model cost $150 and 450M tokens overnight, highlighting the need for careful model selection.
- Infrastructure availability issues with GLM 5.3 Flash due to high popularity forced switching to similar models like DeepSeek V4.1 Flash and Qwen 3.8 Flash.
- Experimentation and R&D with a wide range of models is essential for benchmarking and guiding users toward efficient options.
- Key takeaways include measuring tokens, energy, and cost constantly; budgeting for experimentation; improving prompt selection and multi-agent techniques; and pushing for more efficient models.
- The goal for October is to focus on one or two flash-tier cheap models for the majority of AI inference, measured by cost or energy use rather than token count.