A GPU-Hour Isn't a Commodity If You Need Four of Them
7 hours ago
- Headline GPU rental prices (e.g., $3.93/GPU-hour for H200) mask scarcity for multi-GPU clusters; requesting 4 identical GPUs cuts eligible supply by half and raises price 4%, while 8 GPUs are unavailable.
- Real workloads like training and inference require co-located, identical GPUs with adequate interconnect, not scattered GPU-hours; supply disappears before price spikes, leading to no transaction.
- Compute rations through configuration and availability rather than price; large clusters are often unavailable at any price, driving reliance on bilateral capacity agreements.
- First compute futures based on generalized GPU-hour benchmarks would hedge price but not cluster availability, creating basis risk; future markets may need topology-based contracts for cluster size and interconnect.
- Data from Vast.ai for A100, H100, H200, B200, L40S shows thinner supply for larger clusters; e.g., B200 has no four-GPU offers, H200 no eight-GPU offers, while A100 shows a 114% premium for eight GPUs.
- Caveats: sample is from one marketplace, based on advertised offers not transactions, and small cluster estimates rest on few machines, but availability pattern is robust across reliability thresholds.