K2 Horizon 375B A23B is a sparse Mixture-of-Experts model designed for agentic workflows, tool use, terminal tasks, reasoning, coding, and long-horizon inference. It features 375B total parameters with 23B activated per token, using a 61-layer architecture with 6,144 hidden size, 48 attention heads, and 8 KV heads. The architecture employs 192 routed experts, 8 activated per token, and 1 shared expert, with 1,792-dimensional MoE feed-forward layers and SiLU activation. Supporting a native 512K-token context, it uses BF16 computation for efficient high-capability long-context inference.
K2 Horizon 7B is a medium-sized dense model designed for agentic workflows, coding, reasoning, and long-context language tasks. It features 7B parameters with a 36-layer decoder-only architecture using 4,096 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 12,288-dimensional feed-forward layers, with four layer-normalization groups and a 10M RoPE theta. Supporting a native 512K-token context from midtraining onward, it provides a strong dense baseline for long-context workloads and supports faster inference through Diffusion Adapters, with intermediate checkpoints enabling capability analysis across training.
K2 Horizon 3.7B is a compact dense model designed for agentic workflows, coding, reasoning, and long-context language tasks. It features 3.7B parameters with a 36-layer decoder-only architecture using 2,560 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 10,240-dimensional feed-forward layers, with two layer-normalization groups and a 10M RoPE theta. Supporting a native 512K-token context from midtraining onward, it provides an efficient small-model baseline for long-context workloads while enabling capability analysis through released intermediate training checkpoints.
K2 Horizon MoVA 36B A4B is a sparse Mixture-of-Experts model designed for agentic workflows, reasoning, coding, and efficient long-context inference. It features 36B total parameters with 4B activated per token, using a 48-layer architecture with 2,560 hidden size, 32 attention heads, and 8 KV heads. The architecture combines Mixture-of-Values Attention with 100 routed experts, 8 activated per token, 1 shared expert, and 64 MoVA experts with 4 selected per token. Supporting a native 512K-token context, it uses SiLU-activated 6,144-dimensional feed-forward layers for efficient high-capability inference.
K2 Horizon 32B is a large dense model designed for agentic workflows, coding, reasoning, and long-context language tasks. It features 32B parameters with a 64-layer decoder-only architecture using 5,120 hidden size, 64 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 26,624-dimensional feed-forward layers, with four layer-normalization groups and a 10M RoPE theta. Supporting a native 512K-token context from midtraining onward, it provides a strong dense baseline for demanding long-context workloads while enabling capability analysis through released intermediate training checkpoints.
K2 Horizon 0.9B is a compact dense reasoning model designed for mathematics, coding, science, instruction following, and tool-use tasks. It features 0.9B parameters with a 28-layer decoder-only architecture using 1,536 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 5,120-dimensional feed-forward layers, with YaRN RoPE scaling extending the original 8K context to 128K tokens. Trained through multi-teacher distillation across mathematical, coding, STEM, and instruction-following domains, it provides efficient long-context reasoning for resource-constrained deployments and general-purpose language tasks.
Ling 3.0 Flash VL is a native multimodal Mixture-of-Experts model designed for visual reasoning, agentic workflows, and real-world task execution across images and videos. It features 124B total parameters with 5.5B activated per token, using a 42-layer hybrid backbone with 2,560 hidden size, 32 attention heads, and 32 KV heads. The architecture combines KDA and Gated MLA layers in a 5:1 ratio, with 512 experts activating 8 per token alongside shared experts. Its ViT encoder and two-layer projector integrate visual features, supporting up to 256K-token context for multimodal understanding, reasoning, and action.
Ling 3.0 Flash Fin is a finance-enhanced Mixture-of-Experts model designed for financial research, valuation, spreadsheet workflows, source-grounded analysis, and long-horizon agentic tasks. It features 124B total parameters with 5.1B activated per token, using a 42-layer architecture with 2,560 hidden size, 32 attention heads, and 32 KV heads. The architecture employs a hybrid KDA and attention design, 512 routed experts with 8 activated per token, and 1 shared expert, alongside a native Multi-Token Prediction layer. Supporting a 256K-token context window, it is optimized for multi-document financial reasoning, tool-intensive workflows, and professional research outputs.
DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model designed for long-context reasoning, agentic workflows, and input-heavy workloads across text and images. It features 552B backbone parameters, using a 40-layer Causal Encoder-Decoder architecture with 5,120 hidden size, 64 attention heads, and 1 KV head. The model employs Compressed Sparse Attention 2, 384 routed experts with 6 activated per token, 1 shared expert, and 3 Multi-Token Prediction layers, alongside Engram conditional memory and DSpark speculative decoding. Supporting 1M-token context, it uses a 32-layer vision encoder for multimodal processing.
Granite 4.2 30B is the flagship reasoning model in the Granite 4.2 family, designed for advanced coding, mathematics, tool calling, agentic workflows, and multilingual dialogue. It features 30B parameters with a 64-layer dense transformer using 4,096 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention, RoPE with a 50M theta, and a 32,768-dimensional SwiGLU feed-forward network, with native reasoning and flexible thinking modes for balancing quality and latency. Supporting 128K native context with extension to 512K, it is optimized for complex long-context reasoning and enterprise applications.