Params: 118B
Laguna S 2.1 is a Mixture-of-Experts language model designed for agentic coding, reasoning, and long-horizon workflows. It features 118B total parameters with approximately 8B activated per token, using a 48-layer architecture with 3,072 hidden size and 48 attention heads. The model activates 10 experts per token across 256 routed experts alongside a shared expert, combining hybrid sliding window and full attention with grouped-query attention and per-head gating. Supporting up to a 1M token context window, it is optimized for large-scale coding, tool use, and long-context reasoning.
Text GenerationInstruction Following+7