The Economics of Open-Weight Inference
4 hours ago
- Open-weight models can be cheaper than closed models, with the cheapest costing about one-fifth of comparable closed models, but not at all score thresholds.
- Self-hosting open-weight models on older GPUs like the A100 can be more cost-effective than newer ones like the H100 for certain workloads.
- Five-year rental prices for the A100 retain 80% of one-month prices, suggesting older GPUs retain earning capacity for over a decade.
- Long-running, latency-tolerant workloads (e.g., agents, batch evaluation) can route to cost-efficient older hardware.
- New GPU generations do not automatically make older ones obsolete; economic usefulness depends on workload suitability and cost.
- A100 occupancy rose from 74% to 90% despite increased supply, but open-weight demand role is not definitively established.
- Forward term-price marks are analyst estimates, not executable quotes; occupancy data is from tracked providers, not the entire installed base.
- Open-weight portability allows model deployment on various hardware, enabling use of older GPUs when software and licensing permit.