MiMo V2.6 Pro is an ultra-large omnimodal Mixture-of-Experts model designed for coding agents, computer use, cybersecurity, and long-horizon tasks across text, image, video, and audio. It features 1.02T total parameters with 42B activated, using a 70-layer architecture with 6,144 hidden size, 128 attention heads, and 8 KV heads. The architecture interleaves 60 Sliding Window Attention layers with 10 Global Attention layers, using 384 routed experts with 8 activated per token. Supporting a 1M-token context, it integrates vision and audio encoders with Multi-Token Prediction for demanding agentic workloads.
MiMo V2.6 Flash is a native omnimodal Mixture-of-Experts model designed for coding agents, computer use, cybersecurity, and long-horizon tasks across text, image, video, and audio. It features 309B total parameters with 15B activated, using a 48-layer architecture with 4,096 hidden size, 64 attention heads, and 8 KV heads. The architecture interleaves 39 Sliding Window Attention layers with 9 Global Attention layers, using 256 routed experts with 8 activated per token. Supporting a 1M-token context, it integrates vision and audio encoders with Multi-Token Prediction for efficient agentic inference.
Nex N2.5 Pro is a multimodal agentic model designed for long-horizon tasks, computer use, web browsing, scientific research, knowledge work, and visually grounded workflows. The 397B model uses a 60-layer architecture with 4,096 hidden size, 32 attention heads, and 2 KV heads, combining linear attention with full attention every fourth layer. The architecture employs 512 routed experts, 10 activated per token, and a shared expert, with 1,024-dimensional MoE layers and SiLU activation. Supporting a native 262,144-token context, it incorporates a 27-layer vision encoder for image and video understanding.
Nex N2.5 Mini is a multimodal agentic model designed for long-horizon tasks involving computer use, web browsing, and visually grounded interaction. The 30B model uses a 40-layer sparse Mixture-of-Experts architecture with 256 experts and 8 activated per token, combining linear attention with full attention every fourth layer. It has a 2,048-dimensional hidden size, 16 attention heads, and 2 key-value heads, with one additional next-token prediction layer. Its vision encoder supports image and video inputs, while the model provides a 262K-token context window for extended research, knowledge work, and productivity tasks.
Atria Dawn Preview is an agentic model built on the 744B-parameter GLM-5.2 foundation. It uses a 78-layer sparse Mixture-of-Experts architecture with 256 routed experts and 8 activated per token, combining DeepSeek Sparse Attention (DSA) with dense and sparse MLP layers. The model supports a 1M-token context window and 64 attention heads, with 64 key-value heads and one next-token prediction layer. It targets scientific automation, software creation, research, document generation, cybersecurity, and long-horizon tasks requiring iterative tool use and execution.
Intern S2 397B is a multimodal foundation model designed for scientific intelligence, general reasoning, and long-horizon agents. It uses a 60-layer sparse Mixture-of-Experts architecture with 512 experts and 10 activated per token, alongside 32 attention heads and 2 key-value heads. The model combines linear and full attention, with full attention every fourth layer, and supports a 262K-token context window. Its vision encoder processes image and video inputs, while large-scale reinforcement learning spans scientific domains, general reasoning, and interactive agent environments for advanced scientific problem solving and agentic workflows.
LFM2.5 8B A1B is a compact hybrid Mixture-of-Experts model designed for on-device personal assistants, tool use, instruction following, and agentic workflows. It features 8.3B total parameters with 1.5B active, using 24 layers that combine 18 double-gated convolution blocks with 6 GQA attention blocks. The architecture employs a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and 32 experts with 4 activated per token. Supporting a 128K-token context window, it is trained on 38T tokens for efficient, high-throughput multilingual inference across ten languages and diverse edge devices.
LFM2.5 2.6B is a compact hybrid model designed for tool use, instruction following, multi-step agentic tasks, and efficient on-device deployment. It features 2.69B parameters across 30 layers, combining 22 double-gated short convolution blocks with 8 GQA attention blocks. The architecture uses a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and a 10,752-dimensional SwiGLU feed-forward network. Trained on 34T tokens with agentic reinforcement learning, it supports a 131K-token context window and multilingual text generation across 15 languages for efficient long-context agentic workloads.
LFM2.5 1.2B Thinking is a compact hybrid model designed for general-purpose reasoning and efficient on-device deployment. It features 1.2B parameters across 16 layers, combining 10 double-gated convolution blocks with 6 GQA attention blocks. The architecture uses a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and a 12,288-dimensional SwiGLU feed-forward network. Trained on 28T tokens with extended pre-training and multi-stage reinforcement learning, it supports a 32K-token context window and multilingual text generation across English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
LFM2.5 1.2B Instruct is a compact hybrid model designed for general-purpose language tasks and efficient on-device deployment. It features 1.2B parameters across 16 layers, combining 10 double-gated convolution blocks with 6 GQA attention blocks. The architecture uses a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and a 12,288-dimensional SwiGLU feed-forward network. Supporting extended context through positional scaling, it is optimized for fast edge inference and operates under 1GB of memory. It supports a 32K-token context window and multilingual text generation across eight languages.