Kimi K3 is a frontier-scale native multimodal Mixture-of-Experts model designed for long-horizon reasoning, agentic coding, and large-scale knowledge workflows. It is the world's first open 3T-class model, featuring 2.8T total parameters with 104B activated, built on a 93-layer architecture with Kimi Delta Attention (KDA), Gated MLA, and Stable LatentMoE. The model uses a 7,168 hidden size, 96 attention heads, and 896 routed experts, activating 16 experts per token alongside 2 shared experts. Powered by MoonViT-V2 for native vision understanding and supporting a 1M-token context window, Kimi K3 is optimized for multimodal reasoning, long-context coding, advanced agentic workflows, and frontier-scale research applications
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (moonshotai/Kimi-K3). Released under Kimi-K3. Further reading: Paper · Blog
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure


2,304 GB
Enter your email to get access to this content
Requires the vllm/vllm-openai-rocm:kimi-k3 Docker image or a standard vLLM image version greater than v0.26.0.
Benchmarks measured by Vultr on MI355X · vLLM. Throughput and latency vary with concurrency, input length and engine configuration, so treat these as a comparison baseline, not a service guarantee.