DeepSeek
DeepSeek V3.2 Exp is an experimental Mixture-of-Experts (MoE) large language model developed by DeepSeek AI as an intermediate research step toward its next-generation architecture. It builds upon the V3.1-Terminus design while introducing DeepSeek Sparse Attention (DSA). The model features a 685B parameter architecture with ~37B activated parameters, built on a 61-layer transformer with 128 attention heads and a 7,168 hidden size, utilizing 256 routed experts with 8 experts activated per token. It supports a ~160K token context window using YaRN-based rotary scaling.
DeepSeek
DeepSeek V3.2 is a large Mixture-of-Experts (MoE) language model that balances high computational efficiency with exceptional reasoning and agent capabilities. The model features a 685B-class architecture with ~37B activated parameters, built on a 61-layer transformer with 128 attention heads and a 7,168 hidden size, utilizing 256 routed experts with 8 experts activated per token. It integrates DeepSeek Sparse Attention (DSA) to significantly reduce computational complexity while maintaining long-context performance, supporting up to a ~160K token context window.
DeepSeek
DeepSeek V3.2 Speciale is a reasoning-focused Mixture-of-Experts (MoE) large language model developed for advanced mathematical reasoning, coding, and complex analytical tasks. The model features a 685B parameter architecture with ~37B activated parameters, built on a 61-layer transformer with 128 attention heads and a 7,168 hidden size, utilizing 256 routed experts with 8 experts activated per token. It supports up to a ~160K token context window and integrates DeepSeek Sparse Attention (DSA) to reduce computational complexity while maintaining performance.