Params: 1T
Kimi K2.6 is a native multimodal Mixture-of-Experts model designed for long-horizon coding, agentic workflows, and large-scale autonomous task orchestration, succeeding Kimi K2.5. It features 1T total parameters with ~32B active, 61 layers, 7168 hidden size, and 64 attention heads, activating 8 of 384 experts plus 1 shared expert per token. The model supports a 256K context window with MLA attention and YARN RoPE scaling. It integrates a 400M-parameter MoonViT vision encoder and enables large-scale agent swarms for parallel task execution.
Text GenerationInstruction Following+8