DGX Station GB300 Cluster: Two 400G DACs, and Frontier Models
5 hours ago
- Two GB300 DGX Stations (ASUS and HP) connected via two 400G DAC cables achieved 98% line-rate bandwidth (98GB/s on broadcast, reduce, send/receive) and 1.46 µs RDMA latency.
- For large models like GLM-5.2 (433GB), splitting across two stations in tensor parallel boosted throughput from 139 to 1,525 output tokens per second at 32 streams (11x gain), reaching 3,141 at 256 streams.
- GLM-5.3 showed the highest gain: 188 tokens per second on one station to 5,018 on two (27x), while MiniMax-M3 on long prompts went from stalling at 8 streams to climbing at 256.
- For models fitting in one station (GLM-5.3 Flash, DeepSeek V4.1 Flash), prefill/decode disaggregation doubled or tripled throughput; DeepSeek V4.1 Flash hit 5,248 tokens per second.
- The cluster is easy to set up via NVIDIA's instructions and scales well, making two stations a viable alternative to a GPU server for frontier-class inference without a data center.