Params: 1T
Ring-2.6-1T is a trillion-parameter Mixture-of-Experts reasoning model designed for advanced agentic workflows, complex reasoning, and long-horizon task execution. It features an 80-layer architecture with 8,192 hidden size and 64 attention heads, activating 8 experts per token across 256 routed experts and a shared expert. The model incorporates hybrid attention, QK normalization, Multi-Token Prediction, and YaRN context extension for efficient long-context inference. Supporting up to a 256K token context window, it is optimized for enterprise automation, coding, scientific analysis, and multi-step tool-integrated reasoning.
Text GenerationInstruction Following+7