
InclusionAI
Ling 3.0 Flash VL is a native multimodal Mixture-of-Experts model designed for visual reasoning, agentic workflows, and real-world task execution across images and videos. It features 124B total parameters with 5.5B activated per token, using a 42-layer hybrid backbone with 2,560 hidden size, 32 attention heads, and 32 KV heads. The architecture combines KDA and Gated MLA layers in a 5:1 ratio, with 512 experts activating 8 per token alongside shared experts. Its ViT encoder and two-layer projector integrate visual features, supporting up to 256K-token context for multimodal understanding, reasoning, and action.

InclusionAI
Ling 3.0 Flash Fin is a finance-enhanced Mixture-of-Experts model designed for financial research, valuation, spreadsheet workflows, source-grounded analysis, and long-horizon agentic tasks. It features 124B total parameters with 5.1B activated per token, using a 42-layer architecture with 2,560 hidden size, 32 attention heads, and 32 KV heads. The architecture employs a hybrid KDA and attention design, 512 routed experts with 8 activated per token, and 1 shared expert, alongside a native Multi-Token Prediction layer. Supporting a 256K-token context window, it is optimized for multi-document financial reasoning, tool-intensive workflows, and professional research outputs.

InclusionAI
Ling 3.0 Flash is a 124B-parameter native hybrid-reasoning Mixture-of-Experts model with only 5.1B active parameters, optimized for efficient long-context reasoning and production deployment. Its 42-layer architecture combines 35 Kimi Delta Attention (KDA) layers with 7 gated MLA layers in a 5:1 pattern, alongside 512 routed experts with 8 activated per token and 1 shared expert. Designed for complex agentic workflows, it incorporates training across 10,000+ interactive environments for coding, general, and deep-research agents, with support for context lengths up to 256K.

InclusionAI
Ling 3.0 Tiny is a lightweight hybrid-linear Mixture-of-Experts model designed for efficient reasoning, coding, and agentic workflows on local and resource-constrained hardware. It features 7.9B total parameters with 1.3B activated, using a 24-layer architecture with a 1,536 hidden size and 16 attention heads. The model combines Kimi Delta Attention (KDA) and Multi-Head Latent Attention (MLA) in a 3:1 ratio, with 128 routed experts and 1 shared expert, activating 8 routed experts per token. Supporting a 131K-token context window, it is optimized for low-cost deployment while providing hybrid reasoning and efficient long-context processing for local and edge environments.