Hasty Briefsbeta

Bilingual

DeepSeek 4.1 Flash

5 days ago
  • Introduces the smallest model in a new architecture family with native visual understanding, offering greater capability, faster inference, higher throughput, and scalability.
  • Features an asymmetric Causal Encoder–Decoder architecture with 552B parameters (MoE), using only 8B active for input and 16B for output, reducing costs.
  • New pretraining methods and large-scale RL post-training achieve benchmark results ahead of flagship models like DeepSeek-V4-Pro.
  • KV cache is significantly smaller: 1/4 the HBM and 1/8 the SSD storage, cutting cache-hit costs crucial for agent workloads.
  • V4.1-Flash is live on the DeepSeek API with native multimodal support; older V4-Flash and V4-Flash-Vision-Exp are retired and routed to this new model.
  • V4-Pro is being phased out; from September 14, 2026, all deepseek-v4-pro requests route to V4.1-Flash at its rates until V4.1-Pro launches.
  • Official partners WorkBuddy (CodeBuddy) and OpenCode fully support V4.1-Flash.
  • Lower API prices with peak/off-peak pricing (off-peak rates at 50% of peak), effective September 10, 2026.
  • Open-source community support for inference and additional deployment options are being explored, including large-scale deployments with 2,000 GPUs.