Hasty Briefsbeta

双语

From the creator of Redis; run LLM locally with ds4

4 hours ago
  • DwarfStar 4 (ds4) is a lightweight C inference engine for high-memory Mac, CUDA, and ROCm machines, supporting DeepSeek V4/V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next models.
  • The engine uses asymmetric 2-bit quantization to compress routed experts in mixture-of-experts models while keeping critical shared paths precise, enabling large models to run on local hardware.
  • ds4 provides a CLI (`./ds4`), a local API server (`./ds4-server`), and a persistent agent (`./ds4-agent`), all sharing the same model state and cache.
  • It supports long prefix caching by saving prefixes to SSD and resuming via prompt hash, avoiding full re-prefill after restarts.
  • The project uses custom GGUFs and is validated end-to-end against official model outputs for specific model families, not as a generic GGUF runner.
  • Benchmark results on M5 Max 128 GB show generation speeds of ~34–39 T/s and prefill speeds of ~398–790 T/s depending on context length.
  • The engine exposes OpenAI and Anthropic-style APIs for local coding agents to connect easily.