Params: 753B
GLM 5.2 is an advanced Mixture-of-Experts language model designed for long-horizon reasoning, large-scale coding, and million-token context processing. It features a 78-layer architecture with 6,144 hidden size and 64 attention heads, activating 8 experts per token across 256 routed experts and a shared expert. The model incorporates IndexShare sparse attention to improve long-context efficiency and an enhanced multi-token prediction layer for faster speculative decoding. Supporting up to a 1M token context window, it is optimized for complex reasoning, extended coding workflows, and sustained long-context tasks.
Text GenerationInstruction Following+7