Params: 295B
Hy3 is a Mixture-of-Experts (MoE) language model designed for advanced reasoning, coding, agentic workflows, and long-context processing. It features 295B total parameters with 21B activated, using an 80-layer transformer with 4,096 hidden size and 64 attention heads. The architecture includes 192 routed experts with top-8 routing alongside a shared expert, plus a Multi-Token Prediction layer for faster generation. Supporting up to a 256K context window, it is optimized for reliable tool use, complex multi-turn interactions, and high-throughput production inference.
Text GenerationInstruction Following+7