• Pricing
DashboardContact Sales
  • Blogs
  • Discover
  • Docs
  • Community
⌘K

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

Over 80,000,000 Cloud Servers Launched

Over 80,000,000 Cloud Servers Launched

Cloud ComputeCloud GPUBare MetalFile SystemObject StorageBlock StorageManaged DatabasesCDNServerlessKubernetesContainer RegistryDirect ConnectLoad Balancers
RegionsAdvanced NetworkControl PanelOperating SystemsUpload ISO
Industry CloudOne-Click DeploymentUse Cases
Browse AppsBecome a Vendor
FAQDevelopers / APIsVultr DocsServer StatusBug BountyPromotionsSolution PartnersStart-Up Programs
Our TeamNewsBrand AssetsReferral ProgramCreator ProgramCareersSLALegalVultr Trust CenterContactYour Privacy ChoicesSubprocessorsAccessibility
Contact SalesSign Up

© Vultr 2026 | VULTR is a registered trademark of The Constant Company, LLC.

Terms of ServiceAUPDMCAPrivacy PolicyCookie Policy
  1. Docs
  2. Model Library

Model Library

A curated catalog of flagship open-source AI models with transparent, real-world benchmarks. Every model page includes ready-to-run deployment commands for engines like vLLM and SGLang on Vultr Cloud GPU.

68 families174 variants

Xiaomi MiMo V2.6

Xiaomi

Params: 309B – 1.02T

The MiMo V2.6 family from Xiaomi introduces omnimodal Mixture-of-Experts models in Flash and Pro variants at 309B and 1.02T parameters. Trained through unified reinforcement learning, both support 1M-token context for coding, computer-use, and cybersecurity agents.

Text GenerationInstruction Following+8
Runs onB200+

Vultr VultronCoderAtlas

Vultr

Params: 2.1T

The VultronCoderAtlas family from Vultr introduces a 2.1T-parameter sparse Mixture-of-Experts model derived from Kimi K3. It targets coding, multimodal understanding, tool use, and long-context agentic workflows with efficient expert activation.

Text GenerationInstruction Following+8
Runs onMI325X+

InternLM Atria Dawn

Shanghai AI Lab

Params: 744B

The Atria Dawn family from InternLM introduces an agentic model built on the 744B-parameter GLM-5.2 foundation. It targets scientific automation, software creation, research, document generation, cybersecurity, and long-horizon workflows with iterative tool use, 1M-token context, and sparse MoE architecture.

Text GenerationInstruction Following+7
Runs onB200+

InternLM Intern S2

Shanghai AI Lab

Params: 397B

The Intern S2 family from InternLM introduces a 397B multimodal foundation model designed for scientific intelligence, general reasoning, and long-horizon agents. It combines sparse MoE architecture, hybrid linear and full attention, 262K-token context, and vision capabilities for advanced scientific and agentic workflows.

Text GenerationInstruction Following+8
Runs onB200+

InclusionAI Ling 3.0

InclusionAI

Params: 7.9B – 124B

The Ling 3.0 family from InclusionAI introduces efficient hybrid reasoning model designed for production-scale agentic workflows. It emphasizes inference efficiency, strong reasoning and instruction following, long-context capabilities, and reliable execution across coding, general, and deep-research tasks.

Text GenerationInstruction Following+8
Runs onB200+

DeepSeek V4.1

DeepSeek

Params: 552B

The DeepSeek V4.1 family from DeepSeek AI introduces multimodal Mixture-of-Experts model built for long-context reasoning and agentic workloads. It combines million-token contexts, efficient sparse architectures, controllable reasoning, and native vision-language capabilities.

Text GenerationInstruction Following+8
Runs onB200+

Nex N2.5

Nex AGI

Params: 30B – 397B

The Nex-N2.5 family from Nex AGI introduces next-generation agentic models spanning mini, Pro, and Max variants. It advances computer use, web browsing, visual grounding, and long-horizon task execution, with the Max model built on a 1.6T-parameter MoE foundation.

Text GenerationInstruction Following+8
Runs onB200+

IFM K2 Horizon

IFM

Params: 0.9B – 375B

The IFM K2 Horizon family introduces six open models spanning 0.9B to 375B parameters, unified across reasoning, coding, mathematics, and agentic tasks. Its fully open training lifecycle enables reproducibility, research, and flexible deployment from edge devices to enterprise systems.

Text GenerationInstruction Following+7
Runs onB200+

DeepSeek V4

DeepSeek

Params: 284B – 1.6T

The DeepSeek V4 family from DeepSeek AI introduces large-scale Mixture-of-Experts models with up to trillion-parameter capacity and 1M-token context support. It incorporates hybrid attention, advanced optimization, and domain-specialized training to enable efficient long-context reasoning and strong multi-domain performance.

Text GenerationInstruction Following+7
Runs onB200+

Tencent Hy4 Preview

Tencent

Params: 770B

The Hy4 Preview family from Tencent introduces a frontier-class Mixture-of-Experts model with 770B parameters and 49B active parameters. It emphasizes advanced reasoning, million-token context, sparse attention, and efficient agentic workloads.

Text GenerationInstruction Following+7
Runs onB200+

Z.ai GLM 5.3

Z.ai

Params: 320B – 753B

The GLM-5.3 family from Zhipu AI introduces natively multimodal model combining sparse and linear attention with MoE architecture. It targets efficient long-context reasoning, coding, and advanced agentic workflows with million-token context.

Text GenerationInstruction Following+7
Runs onB200+

Alibaba Qwen 3.8

Alibaba

Params: 27B – 2.4T

The Qwen3.8 family from Alibaba Cloud introduces a frontier-class open model designed for coding, professional work, research, and long-horizon agentic tasks. It combines flexible reasoning control with extended context support up to 1M tokens.

Text GenerationInstruction Following+8
Runs onB200+

IBM Granite 4.2

IBM

Params: 3B – 30B

The Granite 4.2 family from IBM introduces dense reasoning models spanning 3B to 30B parameters, designed for coding, mathematics, tool calling, agentic workflows, multilingual dialogue, and efficient long-context enterprise applications.

Text GenerationInstruction Following+7
Runs onB200+

DeepReinforce.AI Ornith 1.5

DeepReinforce.AI

Params: 9B – 397B

The Ornith 1.5 family from DeepReinforce.AI introduces self-improving foundation models through an end-to-end reinforcement learning loop. It spans 397B and 35B MoE and 9B dense models, combining self-generated tasks, task-specific scaffolds, and solution rollouts to strengthen reasoning, coding, and agentic capabilities across diverse workflows.

Text GenerationInstruction Following+7
Runs onB200+

Liquid AI LFM2.5 VL

Liquid AI

Params: 450M – 3B

The LFM2.5 VL family from Liquid AI introduces compact multimodal models optimized for efficient on-device vision-language inference. It spans multiple sizes supporting image understanding, OCR, grounding, multilingual applications, and function calling.

Text GenerationInstruction Following+7
Runs onB200+

NVIDIA Nemotron v3

NVIDIA

Params: 30B – 550B

The Nemotron 3 family is a series of efficient open language models developed by NVIDIA, built on a hybrid Mixture-of-Experts architecture. The models emphasize high token throughput, long-context reasoning, and cost-efficient deployment for agentic and enterprise AI workloads.

Text GenerationInstruction Following+9
Runs onB200+

Meta Muse Glimmer

Meta

Params: 30B

The Muse Glimmer family from Meta introduces compact multimodal language model optimized for local agentic deployment. It combines reasoning, tool use, failure recovery, and multimodal understanding to support reliable autonomous workflows on consumer hardware.

Text GenerationInstruction Following+8
Runs onB200+

Liquid AI LFM2.5

Liquid AI

Params: 230M – 8B

The LFM2.5 family from Liquid AI introduces compact hybrid models optimized for efficient on-device language, reasoning, tool use, and agentic workloads. It spans 230M to 8.3B parameters, combining convolution and GQA attention architectures with extended context support up to 131K tokens.

Text GenerationInstruction Following+7
Runs onB200+

Thinking Machines Inkling

Thinking Machines

Params: 276B – 975B

The Inkling family from Thinking Machines introduces open-weight multimodal language model designed for reasoning, coding, and agentic applications. It combines native text, image, audio, and video understanding with efficient Mixture-of-Experts architecture to support conversational, tool-use, and long-context workflows.

Text GenerationInstruction Following+8
Runs onB200+

Moonshot AI Kimi K3

Moonshot AI

Params: 2.8T

The Kimi K3 family from Moonshot AI introduces next-generation multimodal Mixture-of-Experts model built for frontier-scale reasoning, coding, and agentic workflows. As the world's first open 3T-class model family, it combines native multimodal understanding, million-token context, and advanced attention architectures for efficient long-horizon performance.

Text GenerationInstruction Following+8
Runs onMI355X+