Params: 357B
GLM 4.6 is a Mixture-of-Experts (MoE) large language model developed by Z.ai designed for advanced coding, reasoning, and agentic workflows. The model features a 357B parameter architecture built on a 92-layer transformer with 96 attention heads and a 5,120 hidden size, utilizing 160 routed experts with 8 experts activated per token for efficient scaling. It supports a ~200K token context window, enabling long-horizon reasoning and complex multi-step agent interactions.
Text GenerationInstruction Following+7