Params: 744B
GLM 5 is a large Mixture-of-Experts (MoE) language model developed by Z.ai for complex reasoning, coding, and long-horizon agentic workflows. The model features a 744B parameter architecture with 40B activated parameters, built on a 78-layer transformer with 64 attention heads and a 6,144 hidden size, utilizing 256 routed experts with 8 experts activated per token. It integrates DeepSeek Sparse Attention (DSA) to reduce deployment cost while maintaining performance, and supports a ~202K token context window for large-scale multi-step reasoning tasks.
Text GenerationInstruction Following+7