
Z.ai
GLM 5.3 is an advanced Mixture-of-Experts model designed for long-horizon reasoning and large-scale coding, delivering substantially improved performance over GLM-5.2 through post-training. Its architecture features 78 layers with hybrid linear and DeepSeek sparse attention, 64 attention heads, 64 KV heads, and 256 routed experts with 8 activated per token. The model incorporates IndexShare sparse attention to improve long-context efficiency and an enhanced multi-token prediction layer for faster speculative decoding. Supporting up to a 1M token context window, it is optimized for complex reasoning, extended coding workflows, and sustained long-context tasks.

Z.ai
GLM 5.3 Flash is a natively multimodal model designed for efficient, high-performance agentic workloads, coding, and long-context applications. It has 320B total parameters with 18B active parameters, delivering better performance than GLM 5.2 while substantially reducing serving costs. The model uses a hybrid sparse architecture combining linear attention with DeepSeek-style sparse attention, alongside 288 routed experts, 8 activated experts per token, and Manifold-Constrained Hyper-Connections (mHC) for improved scaling efficiency. It supports up to 1M tokens context window and multimodal capabilities for demanding agentic workloads.