
InclusionAI
Ling 3.0 Flash VL is a native multimodal Mixture-of-Experts model designed for visual reasoning, agentic workflows, and real-world task execution across images and videos. It features 124B total parameters with 5.5B activated per token, using a 42-layer hybrid backbone with 2,560 hidden size, 32 attention heads, and 32 KV heads. The architecture combines KDA and Gated MLA layers in a 5:1 ratio, with 512 experts activating 8 per token alongside shared experts. Its ViT encoder and two-layer projector integrate visual features, supporting up to 256K-token context for multimodal understanding, reasoning, and action.
DeepSeek
DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model designed for long-context reasoning, agentic workflows, and input-heavy workloads across text and images. It features 552B backbone parameters, using a 40-layer Causal Encoder-Decoder architecture with 5,120 hidden size, 64 attention heads, and 1 KV head. The model employs Compressed Sparse Attention 2, 384 routed experts with 6 activated per token, 1 shared expert, and 3 Multi-Token Prediction layers, alongside Engram conditional memory and DSpark speculative decoding. Supporting 1M-token context, it uses a 32-layer vision encoder for multimodal processing.
DeepSeek
DeepSeek V4 Flash Vision Exp is a multimodal Mixture-of-Experts model designed for visual understanding and multimodal agentic workflows while maintaining strong text-only agent performance. It features 284B total parameters with 13B activated, using a 43-layer architecture with 4,096 hidden size, 64 attention heads, and 1 KV head. The model employs 256 routed experts with 6 activated per token alongside a shared expert, DeepSeek sparse attention, and 3 Multi-Token Prediction layers. Supporting up to a 1M-token context, it integrates a 32-layer vision encoder for enhanced multimodal reasoning and visual processing.

Tencent
Hy4 preview is a flagship Mixture-of-Experts model designed for advanced reasoning, coding, agentic workflows, and efficient long-context processing. It features 770B total parameters with 49B activated per token, using 78 layers with 6,144 hidden size and 64 attention heads. The architecture combines Gated DeepSeek Sparse Attention with IndexCache, 256 routed experts and 1 shared expert, activating 8 routed experts per token, while iHC expands residual information flow across 4 streams. It includes a native Multi-Token Prediction layer for speculative decoding and supports up to a 1M-token context window.

Z.ai
GLM 5.3 is an advanced Mixture-of-Experts model designed for long-horizon reasoning and large-scale coding, delivering substantially improved performance over GLM-5.2 through post-training. Its architecture features 78 layers with hybrid linear and DeepSeek sparse attention, 64 attention heads, 64 KV heads, and 256 routed experts with 8 activated per token. The model incorporates IndexShare sparse attention to improve long-context efficiency and an enhanced multi-token prediction layer for faster speculative decoding. Supporting up to a 1M token context window, it is optimized for complex reasoning, extended coding workflows, and sustained long-context tasks.

Alibaba
Qwen3.8 Flash Next is a 125B parameter hybrid model with 6B activated parameters, along with 51B n-gram embedding and 4B MTP parameters, designed for efficient long-context and agentic workloads. Its architecture combines Gated DeltaNet with Qwen Sparse Attention, 512 experts with 10 routed and 1 shared expert, Gated Residuals, and n-gram embeddings for efficient scaling and inference. It supports 262K-token context, extensible to 1M tokens, with native vision and video understanding. The model targets efficient deployment while reducing long-context latency and memory demands.
IBM
Granite 4.2 30B is the flagship reasoning model in the Granite 4.2 family, designed for advanced coding, mathematics, tool calling, agentic workflows, and multilingual dialogue. It features 30B parameters with a 64-layer dense transformer using 4,096 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention, RoPE with a 50M theta, and a 32,768-dimensional SwiGLU feed-forward network, with native reasoning and flexible thinking modes for balancing quality and latency. Supporting 128K native context with extension to 512K, it is optimized for complex long-context reasoning and enterprise applications.

DeepReinforce.AI
Ornith 1.5 35B A3B is the mid-size Mixture-of-Experts member of the Ornith-1.5 family, designed for coding and agentic workloads with only ~3B parameters activated per token. Its end-to-end self-improvement approach jointly optimizes task generation, scaffold construction, and solution rollouts through reinforcement learning. The Qwen3.5-based architecture features 40 layers with hybrid linear and full attention, 16 attention heads, 2 KV heads, and 256 experts with 8 activated per token. The model supports a 262,144 token context window.

Liquid AI
LFM2.5 VL 3B is a compact multimodal language model designed for efficient on-device vision-language tasks, combining the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It uses a 30-layer hybrid architecture with 2,048 hidden size, 32 attention heads, and 8 key-value heads, alternating short convolution and full-attention layers for efficient sequence processing. Supporting a 32K-token context window, it uses native-resolution image processing with 512×512 patches and thumbnail processing. The model is optimized for image grounding, object detection, full-page OCR, layout understanding, multilingual applications, and efficient on-device multimodal inference.

NVIDIA
NVIDIA Nemotron 3.5 Lightning 30B A3B is a 30B-parameter hybrid Mixture-of-Experts model with 3B active parameters, combining Mamba-2, MoE, and Transformer attention layers. It features 52 layers, 128 routed experts with 6 activated per token, 32 attention heads, and a 2-head KV configuration, supporting up to 1M-token context, with multi-token prediction and reinforcement learning for reasoning, coding, tool use, and agentic workflows.

Meta
Muse Glimmer 30B is a multimodal dense causal language model designed for autonomous agentic tasks, combining multi-step reasoning, reliable tool use, failure recovery, and visual understanding. It has 30B parameters, including a dedicated 1.8B-parameter perception encoder, with a 52-layer language model using a 6,656 hidden size and 32 attention heads. The language model uses a repeating 3:1 local-to-global attention pattern with a 2,048-token sliding window and GQA with 32 query and 2 KV heads. It supports text and image inputs with a 131K+ context window and is optimized for local deployment and long-horizon agent workflows.

Thinking Machines
Inkling Small is a multimodal Mixture-of-Experts model with 276B total parameters and 12B active parameters, designed for conversational, coding, retrieval, and agentic applications. It features a 42-layer decoder-only architecture with 32 attention heads, 256 routed experts activating 6 per token, and 2 shared experts. Supporting text, image, and audio inputs, it uses hybrid local and global attention with a 1M-token context window and a 512-token sliding window, while supporting BF16 and NVFP4 numerics.

Moonshot AI
Kimi K3 is a frontier-scale native multimodal Mixture-of-Experts model designed for long-horizon reasoning, agentic coding, and large-scale knowledge workflows. It is the world's first open 3T-class model, featuring 2.8T total parameters with 104B activated, built on a 93-layer architecture with Kimi Delta Attention (KDA), Gated MLA, and Stable LatentMoE. The model uses a 7,168 hidden size, 96 attention heads, and 896 routed experts, activating 16 experts per token alongside 2 shared experts. Powered by MoonViT-V2 for native vision understanding and supporting a 1M-token context window, Kimi K3 is optimized for multimodal reasoning, long-context coding, advanced agentic workflows, and frontier-scale research applications

Mindlab Research
Macaron V1 Venti is a 753B parameter Mixture of LoRA (MoL) flagship model built on a 744B GLM-5.2 base with four 1B-parameter specialists for chat, personal-agent tasks, coding, and Generative UI. It uses a 78-layer architecture with a 6,144 hidden size and 64 attention heads, incorporating sparse MoE layers with 256 routed experts and 1 shared expert, activating 8 routed experts per token. The model integrates IndexShare sparse attention and multi-token prediction, while supporting a 1M-token context window for long-horizon agentic workflows, coding, tool use, and Generative UI.

Poolside
Laguna S 2.1 is a Mixture-of-Experts language model designed for agentic coding, reasoning, and long-horizon workflows. It features 118B total parameters with approximately 8B activated per token, using a 48-layer architecture with 3,072 hidden size and 48 attention heads. The model activates 10 experts per token across 256 routed experts alongside a shared expert, combining hybrid sliding window and full attention with grouped-query attention and per-head gating. Supporting up to a 1M token context window, it is optimized for large-scale coding, tool use, and long-context reasoning.

Tencent
Hy3 is a Mixture-of-Experts (MoE) language model designed for advanced reasoning, coding, agentic workflows, and long-context processing. It features 295B total parameters with 21B activated, using an 80-layer transformer with 4,096 hidden size and 64 attention heads. The architecture includes 192 routed experts with top-8 routing alongside a shared expert, plus a Multi-Token Prediction layer for faster generation. Supporting up to a 256K context window, it is optimized for reliable tool use, complex multi-turn interactions, and high-throughput production inference.

DeepReinforce.AI
Ornith 1.0 397B is a 397B parameter sparse Mixture-of-Experts model designed for high-performance agentic coding and software-engineering workflows. Built on Qwen3.5, it uses a 60-layer hybrid architecture combining linear and full attention, with a 4,096 hidden size, 32 attention heads, 2 KV heads, and 512 experts activating 10 per token. Approximately 17B parameters are active per token, with a 262K-token context window and multimodal support. Its self-improving RL framework jointly optimizes solution rollouts and scaffolds, enabling stronger search trajectories and higher-quality coding solutions.

Alibaba
Qwen AgentWorld 35B A3B is a native language world model designed to simulate agent environments across MCP, Search, Terminal, SWE, Android, Web, and OS interaction domains. It features 35B total parameters with 3B activated, using a 40-layer hybrid architecture with 2,048 hidden size and 16 attention heads. The model combines Gated DeltaNet and Gated Attention layers in a 3:1 pattern with 256 routed experts, activating 8 experts per token alongside a shared expert. Trained through CPT, SFT, and RL, it supports native world modeling and a 262K-token context window.
Vultr
VultronRetrieverPrime Qwen3.5 8B is a multimodal late-interaction retrieval model designed for visual document search and multilingual RAG across PDFs, scans, slides, and reports. It features 8B parameters, using a 32-layer hybrid GatedDeltaNet and full-attention architecture with 4,096 hidden size, 16 attention heads, and 320-dimensional multi-vector embeddings with MaxSim scoring. Supporting up to 262K context and 1,792 visual tokens, it delivers state-of-the-art retrieval accuracy while maintaining a compact index and efficient large-scale serving.

Z.ai
GLM 5.2 is an advanced Mixture-of-Experts language model designed for long-horizon reasoning, large-scale coding, and million-token context processing. It features a 78-layer architecture with 6,144 hidden size and 64 attention heads, activating 8 experts per token across 256 routed experts and a shared expert. The model incorporates IndexShare sparse attention to improve long-context efficiency and an enhanced multi-token prediction layer for faster speculative decoding. Supporting up to a 1M token context window, it is optimized for complex reasoning, extended coding workflows, and sustained long-context tasks.