
Alibaba
Qwen3.8 2.4T A95B is a frontier-scale Mixture-of-Experts model designed for advanced reasoning, coding, professional work, research, and long-horizon agentic tasks. It features 2.4T total parameters with 95B activated, using a 92-layer hybrid architecture with an 8,192 hidden size and 64 attention heads. The model combines Gated DeltaNet linear attention with Gated Attention, using 512 experts with 10 routed experts and 1 shared expert activated per token. Supporting a native 262K-token context window extensible to approximately 1.01M tokens, it is optimized for complex reasoning, autonomous agent execution, coding, and sustained long-context workloads.

Alibaba
Qwen3.8 Flash Next is a 125B parameter hybrid model with 6B activated parameters, along with 51B n-gram embedding and 4B MTP parameters, designed for efficient long-context and agentic workloads. Its architecture combines Gated DeltaNet with Qwen Sparse Attention, 512 experts with 10 routed and 1 shared expert, Gated Residuals, and n-gram embeddings for efficient scaling and inference. It supports 262K-token context, extensible to 1M tokens, with native vision and video understanding. The model targets efficient deployment while reducing long-context latency and memory demands.

Alibaba
Qwen3.8 27B is a compact dense vision-language model designed for coding, professional work, research, and long-horizon agentic tasks. It features 27B parameters across 64 layers, with a 5,120 hidden size and 24 attention heads, using a hybrid architecture that combines Gated DeltaNet linear attention with full attention in a 3:1 pattern. The model includes a native vision encoder for image and video understanding, supports flexible reasoning control, and integrates Multi-Token Prediction. It provides a 262K-token native context window, extensible to 1M tokens.