Hasty Briefsbeta

Bilingual

Hetzner is working on LLM Inference

3 hours ago
  • Hetzner is experimenting with LLM inference via an OpenAI-compatible API, but it's still in early stages with no billing, SLA, or production guarantee.
  • The only model currently available is Qwen/Qwen3.6-35B-A3B-FP8, a 35B-parameter Mixture-of-Experts model with 3B active parameters, supporting text and images.
  • Initial tests show fast performance: 153 ms median time to first token and 224 output tokens per second, though these are not representative of production loads.
  • The product's potential lies in Hetzner's ability to leverage its low-cost hardware operations and GPU utilization, but the current GPU lineup lacks large-scale multi-GPU systems needed for bigger models.
  • The experiment's future depends on whether Hetzner expands to larger GPU clusters and a broader model catalog, which could make it a serious competitor in inference services.