Hasty Briefsbeta

Bilingual

Tokens Too Cheap to Meter

7 hours ago
  • The cost of using AI is decreasing by orders of magnitude annually, enabling integration of LLMs into computing infrastructure within 1-2 years and local frontier-quality models within 3-6 years.
  • Improvements in GPUs, model efficiency (e.g., Mixture-of-Experts), and inference engines (e.g., vLLM, NVIDIA, Intel) are driving a roughly 2.5 orders of magnitude decrease in token cost per year.
  • New architectures like Mamba-Transformer hybrids reduce RAM requirements for local AI by 5x or more, allowing larger models to run on commodity hardware.
  • Specialized models like Jev and Laya offer extreme cost reductions (e.g., $42 per billion tokens for Jev) for specific tasks, making AI cheaper than many traditional computing tools.
  • Supply-side Jevons Paradox suggests cheaper AI leads to increased compute usage, benefiting hyperscalers and inference-optimized hardware providers.
  • Demand-side Jevons Paradox raises questions about future use cases, including cybersecurity shifts, software commoditization (operations over code), and increased QA/UI roles.
  • The fourth option of using LLMs to build software creates new competition and enables personalized, malleable software for non-programmers.