DeepSeek
DeepSeek V4 Pro 0813 is the official release of DeepSeek-V4-Pro, superseding the Preview version with enhanced agentic capabilities and improved performance, particularly in production environments. It features 1.6T total parameters with approximately 49B activated, using 384 routed experts with 6 selected per token across a 61-layer architecture with 7,168 hidden size and 128 attention heads. The model incorporates DeepSeek’s compressed long-context architecture with sparse attention, multi-token prediction, and YaRN RoPE scaling to support up to a 1M-token context window. It also includes the DSpark speculative decoding module, with dedicated target layers and a 5-token block size, for faster generation.
DeepSeek
DeepSeek V4 Pro is an ultra-large Mixture-of-Experts model designed for high-performance long-context reasoning and large-scale deployment. It features 1.6T total parameters with approximately 49B activated, using 384 routed experts with 6 selected per token across a 61-layer architecture with 7,168 hidden size and 128 attention heads. Built with hybrid compressed attention mechanisms and manifold-constrained hyper-connections, it supports up to a 1M token context window. With FP4 and FP8 mixed precision and advanced optimization techniques, it delivers strong efficiency, stability, and agentic reasoning performance across complex workloads.
DeepSeek
DeepSeek V4 Flash Vision Exp is a multimodal Mixture-of-Experts model designed for visual understanding and multimodal agentic workflows while maintaining strong text-only agent performance. It features 284B total parameters with 13B activated, using a 43-layer architecture with 4,096 hidden size, 64 attention heads, and 1 KV head. The model employs 256 routed experts with 6 activated per token alongside a shared expert, DeepSeek sparse attention, and 3 Multi-Token Prediction layers. Supporting up to a 1M-token context, it integrates a 32-layer vision encoder for enhanced multimodal reasoning and visual processing.
DeepSeek
DeepSeek V4 Flash 0731 is a large-scale Mixture-of-Experts model optimized for ultra-long-context reasoning, efficient inference, and enhanced agentic capabilities, with 284B total parameters and approximately 13B activated. It uses 256 routed experts, activating 6 per token, across 43 layers with a 4,096 hidden size and 64 attention heads. Its architecture combines CSA and HCA attention with manifold-constrained hyper-connections and supports a 1M-token context window. It features FP4 expert weights, FP8 quantization, and a speculative decoding module for faster generation.
DeepSeek
DeepSeek V4 Flash is a large-scale Mixture-of-Experts model optimized for ultra-long context reasoning and efficient inference. It features 284B total parameters with approximately 13B activated, using 256 routed experts with 6 selected per token across a 43-layer architecture with 4,096 hidden size and 64 attention heads. Built with hybrid CSA + HCA attention and manifold-constrained hyper-connections, it supports up to a 1M token context window. With FP4 and FP8 mixed precision and strong agentic tool-calling capabilities, it is designed for scalable, high-efficiency reasoning and multi-domain workloads.