
NVIDIA
Cosmos Reason2 8B is a multimodal vision-language model optimized for physical AI and robotics. It contains 8B parameters, built on a ViT vision encoder and a 36-layer dense transformer with 32 attention heads and 4096 hidden size. Post-trained from Qwen3 VL 8B Instruct, it supports up to 256K token context and spatio-temporal reasoning for robotics, video analytics, and data curation. The model enables object detection with 2D/3D localization, common-sense planning, and embodied reasoning, serving as the foundation for AI agents to interpret, plan, and act in complex real-world environments.

NVIDIA
Cosmos Reason2 2B is a multimodal vision-language model optimized for physical AI and robotics. It contains 2.4B parameters, built on a ViT vision encoder and a 28-layer dense transformer with 16 attention heads and 2,048 hidden size. Post-trained from Qwen3 VL 2B Instruct, it supports up to 256K token context and spatio-temporal reasoning for robotics, video analytics, and data curation. The model enables object detection with 2D/3D localization, common-sense planning, and embodied reasoning, serving as the foundation for AI agents to interpret, plan, and act in complex real-world environments.