Intern S2 397B is a multimodal foundation model designed for scientific intelligence, general reasoning, and long-horizon agents. It uses a 60-layer sparse Mixture-of-Experts architecture with 512 experts and 10 activated per token, alongside 32 attention heads and 2 key-value heads. The model combines linear and full attention, with full attention every fourth layer, and supports a 262K-token context window. Its vision encoder processes image and video inputs, while large-scale reinforcement learning spans scientific domains, general reasoning, and interactive agent environments for advanced scientific problem solving and agentic workflows.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (internlm/Intern-S2-397B). Released under Apache 2.0. Further reading: Blog
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure
1,440 GB
Enter your email to get access to this content
Benchmarks measured by Vultr on B200 · vLLM. Throughput and latency vary with concurrency, input length and engine configuration, so treat these as a comparison baseline, not a service guarantee.