Kimi K2.5 is a native multimodal Mixture-of-Experts (MoE) large language model designed for advanced coding, vision reasoning, and autonomous agentic workflows. The model features a 1T parameter architecture with 32B activated parameters, 61 layers, 64 attention heads, and 384 experts (8 experts per token). It supports up to a 256K token context window and integrates a 400M-parameter MoonViT vision encoder for cross-modal understanding. The model employs native INT4 quantization with Quantization-Aware Training (QAT) to reduce inference latency and memory usage while maintaining strong performance.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (moonshotai/Kimi-K2.5). Released under Modified MIT. Further reading: Paper · Blog
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure


2,048 GB
1,440 GB
Enter your email to get access to this content
Benchmarks measured by Vultr on MI325X · vLLM. Throughput and latency vary with concurrency, input length and engine configuration, so treat these as a comparison baseline, not a service guarantee.