Cosmos Reason2 8B is a multimodal vision-language model optimized for physical AI and robotics. It contains 8B parameters, built on a ViT vision encoder and a 36-layer dense transformer with 32 attention heads and 4096 hidden size. Post-trained from Qwen3 VL 8B Instruct, it supports up to 256K token context and spatio-temporal reasoning for robotics, video analytics, and data curation. The model enables object detection with 2D/3D localization, common-sense planning, and embodied reasoning, serving as the foundation for AI agents to interpret, plan, and act in complex real-world environments.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (nvidia/Cosmos-Reason2-8B). Released under nvidia-open-model-license. Further reading: Blog
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure
1,440 GB
Enter your email to get access to this content
Nvidia Cosmos Reason 2 is a gated model; ensure you have been granted access on Hugging Face and have authenticated your environment using a valid HF_TOKEN.
Benchmarks measured by Vultr on B200 · vLLM. Throughput and latency vary with concurrency, input length and engine configuration, so treat these as a comparison baseline, not a service guarantee.