Hasty Briefsbeta

Bilingual

The Economics of Open-Weight Inference

4 hours ago
  • Open-weight models can be cheaper than closed models, with the cheapest costing about one-fifth of comparable closed models, but not at all score thresholds.
  • Self-hosting open-weight models on older GPUs like the A100 can be more cost-effective than newer ones like the H100 for certain workloads.
  • Five-year rental prices for the A100 retain 80% of one-month prices, suggesting older GPUs retain earning capacity for over a decade.
  • Long-running, latency-tolerant workloads (e.g., agents, batch evaluation) can route to cost-efficient older hardware.
  • New GPU generations do not automatically make older ones obsolete; economic usefulness depends on workload suitability and cost.
  • A100 occupancy rose from 74% to 90% despite increased supply, but open-weight demand role is not definitively established.
  • Forward term-price marks are analyst estimates, not executable quotes; occupancy data is from tracked providers, not the entire installed base.
  • Open-weight portability allows model deployment on various hardware, enabling use of older GPUs when software and licensing permit.