DeepSeek-v4.1-Exp
21 days ago
- DeepSeek-V4.1-Flash is a multimodal MoE model with 552B backbone parameters, supporting up to 1M token context and native image-text processing.
- It uses a Causal Encoder-Decoder (CED) architecture with compressed KV cache (CSA2), achieving 890 bytes per token—1/4 of V4-Flash—and activating only 8B/16B parameters per token.
- The model is trained from scratch on 45T tokens and includes a vision encoder (DeepSeek-ViT) for multimodal capabilities.
- Post-training follows SFT→RL→OPD with automated agent task synthesis, and features a controllable reasoning effort setting (1-100).
- Evaluation shows competitive performance on agentic, reasoning, code, math, and multimodal benchmarks, outperforming many frontier models.
- Usage is supported via libraries (Transformers, vLLM, SGLang), Docker, notebooks (Colab, Kaggle), and local apps with detailed instructions provided.
- Prompt encoding is supplied as a Python reference and a Rust library (deepseek-recipe) for production use.
- The model and code are released under the MIT License.