<antirez>
5 hours ago
- DwarfStar 4 (DS4) became popular quickly due to the release of a quasi-frontier model that is large and fast enough for local inference, combined with an asymmetric 2/8 bit quant setup that works on 96-128GB RAM.
- The project leverages the experience of the local AI movement and GPT 5.5, allowing rapid development (one week) with proper LLM interaction skills.
- The author worked 14 hours/day on average during the last week, similar to the early Redis days, but normal average is 4-6 hours.
- DS4 is not a one-time project; it will evolve to use the best open weights model that is practically fast on high-end Mac or 'GPU in a box' setups like DGX Spark.
- The next contender is expected to be DeepSeek v4 Flash itself, possibly with a new checkpoint and coding-tuned version, plus other domain-specific variants (coding, legal, medical).
- This is the first time the author uses a local model for serious tasks normally done with Claude/GPT, and vector steering enables more freedom with the LLM.
- Future focus includes quality benchmarks, adding a coding agent, a home hardware CI setup for long-term quality, more ports, and distributed inference (serial and parallel).
- The author emphasizes that AI is too critical to be just a provided service.