Params: 552B
DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model designed for long-context reasoning, agentic workflows, and input-heavy workloads across text and images. It features 552B backbone parameters, using a 40-layer Causal Encoder-Decoder architecture with 5,120 hidden size, 64 attention heads, and 1 KV head. The model employs Compressed Sparse Attention 2, 384 routed experts with 6 activated per token, 1 shared expert, and 3 Multi-Token Prediction layers, alongside Engram conditional memory and DSpark speculative decoding. Supporting 1M-token context, it uses a 32-layer vision encoder for multimodal processing.
Text GenerationInstruction Following+8