IFM
K2 Horizon 375B A23B is a sparse Mixture-of-Experts model designed for agentic workflows, tool use, terminal tasks, reasoning, coding, and long-horizon inference. It features 375B total parameters with 23B activated per token, using a 61-layer architecture with 6,144 hidden size, 48 attention heads, and 8 KV heads. The architecture employs 192 routed experts, 8 activated per token, and 1 shared expert, with 1,792-dimensional MoE feed-forward layers and SiLU activation. Supporting a native 512K-token context, it uses BF16 computation for efficient high-capability long-context inference.
IFM
K2 Horizon MoVA 36B A4B is a sparse Mixture-of-Experts model designed for agentic workflows, reasoning, coding, and efficient long-context inference. It features 36B total parameters with 4B activated per token, using a 48-layer architecture with 2,560 hidden size, 32 attention heads, and 8 KV heads. The architecture combines Mixture-of-Values Attention with 100 routed experts, 8 activated per token, 1 shared expert, and 64 MoVA experts with 4 selected per token. Supporting a native 512K-token context, it uses SiLU-activated 6,144-dimensional feed-forward layers for efficient high-capability inference.
IFM
K2 Horizon 32B is a large dense model designed for agentic workflows, coding, reasoning, and long-context language tasks. It features 32B parameters with a 64-layer decoder-only architecture using 5,120 hidden size, 64 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 26,624-dimensional feed-forward layers, with four layer-normalization groups and a 10M RoPE theta. Supporting a native 512K-token context from midtraining onward, it provides a strong dense baseline for demanding long-context workloads while enabling capability analysis through released intermediate training checkpoints.
IFM
K2 Horizon 7B is a medium-sized dense model designed for agentic workflows, coding, reasoning, and long-context language tasks. It features 7B parameters with a 36-layer decoder-only architecture using 4,096 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 12,288-dimensional feed-forward layers, with four layer-normalization groups and a 10M RoPE theta. Supporting a native 512K-token context from midtraining onward, it provides a strong dense baseline for long-context workloads and supports faster inference through Diffusion Adapters, with intermediate checkpoints enabling capability analysis across training.
IFM
K2 Horizon 3.7B is a compact dense model designed for agentic workflows, coding, reasoning, and long-context language tasks. It features 3.7B parameters with a 36-layer decoder-only architecture using 2,560 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 10,240-dimensional feed-forward layers, with two layer-normalization groups and a 10M RoPE theta. Supporting a native 512K-token context from midtraining onward, it provides an efficient small-model baseline for long-context workloads while enabling capability analysis through released intermediate training checkpoints.
IFM
K2 Horizon 0.9B is a compact dense reasoning model designed for mathematics, coding, science, instruction following, and tool-use tasks. It features 0.9B parameters with a 28-layer decoder-only architecture using 1,536 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention and SiLU-activated 5,120-dimensional feed-forward layers, with YaRN RoPE scaling extending the original 8K context to 128K tokens. Trained through multi-teacher distillation across mathematical, coding, STEM, and instruction-following domains, it provides efficient long-context reasoning for resource-constrained deployments and general-purpose language tasks.