Params: 2.1T
VultronCoderAtlas is a 2.1T-parameter sparse Mixture-of-Experts model with approximately 107B active parameters, designed for repository-scale coding, agentic tool use, long-context technical analysis, and multimodal workloads. Derived from Kimi K3, it uses 93 transformer blocks comprising 1 dense and 92 MoE layers, with 672 routed experts, 16 activated per token, and 2 shared experts. The architecture combines 23 MLA and 69 KDA blocks with a 7,168-dimensional hidden size and 96 attention heads. Supporting text and image inputs across a 1M-token context, it applies REAP expert pruning while retaining the native Kimi K3 vision pathway.
Text GenerationInstruction Following+8