• Pricing
DashboardContact Sales
  • Blogs
  • Discover
  • Docs
  • Community
⌘K

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

 

 

  

Runs on

Over 80,000,000 Cloud Servers Launched

Over 80,000,000 Cloud Servers Launched

Cloud ComputeCloud GPUBare MetalFile SystemObject StorageBlock StorageManaged DatabasesCDNServerlessKubernetesContainer RegistryDirect ConnectLoad Balancers
RegionsAdvanced NetworkControl PanelOperating SystemsUpload ISO
Industry CloudOne-Click DeploymentUse Cases
Browse AppsBecome a Vendor
FAQDevelopers / APIsVultr DocsServer StatusBug BountyPromotionsSolution PartnersStart-Up Programs
Our TeamNewsBrand AssetsReferral ProgramCreator ProgramCareersSLALegalVultr Trust CenterContactYour Privacy ChoicesSubprocessorsAccessibility
Contact SalesSign Up

© Vultr 2026 | VULTR is a registered trademark of The Constant Company, LLC.

Terms of ServiceAUPDMCAPrivacy PolicyCookie Policy
  1. Docs
  2. Model Library

Model Library

A curated catalog of flagship open-source AI models with transparent, real-world benchmarks. Every model page includes ready-to-run deployment commands for engines like vLLM and SGLang on Vultr Cloud GPU.

61 families155 variants

InclusionAI Ling 3.0

InclusionAI

Params: 7.9B – 124B

Ling 3.0 Flash VL is a native multimodal Mixture-of-Experts model designed for visual reasoning, agentic workflows, and real-world task execution across images and videos. It features 124B total parameters with 5.5B activated per token, using a 42-layer hybrid backbone with 2,560 hidden size, 32 attention heads, and 32 KV heads. The architecture combines KDA and Gated MLA layers in a 5:1 ratio, with 512 experts activating 8 per token alongside shared experts. Its ViT encoder and two-layer projector integrate visual features, supporting up to 256K-token context for multimodal understanding, reasoning, and action.

Text GenerationInstruction Following+8
Runs onB200+

DeepSeek V4.1

DeepSeek

Params: 552B

DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model designed for long-context reasoning, agentic workflows, and input-heavy workloads across text and images. It features 552B backbone parameters, using a 40-layer Causal Encoder-Decoder architecture with 5,120 hidden size, 64 attention heads, and 1 KV head. The model employs Compressed Sparse Attention 2, 384 routed experts with 6 activated per token, 1 shared expert, and 3 Multi-Token Prediction layers, alongside Engram conditional memory and DSpark speculative decoding. Supporting 1M-token context, it uses a 32-layer vision encoder for multimodal processing.

Text GenerationInstruction Following+8
Runs onB200+

DeepSeek V4

DeepSeek

Params: 284B – 1.6T

DeepSeek V4 Flash Vision Exp is a multimodal Mixture-of-Experts model designed for visual understanding and multimodal agentic workflows while maintaining strong text-only agent performance. It features 284B total parameters with 13B activated, using a 43-layer architecture with 4,096 hidden size, 64 attention heads, and 1 KV head. The model employs 256 routed experts with 6 activated per token alongside a shared expert, DeepSeek sparse attention, and 3 Multi-Token Prediction layers. Supporting up to a 1M-token context, it integrates a 32-layer vision encoder for enhanced multimodal reasoning and visual processing.

Text GenerationInstruction Following+7
Runs onB200+

Tencent Hy4 Preview

Tencent

Params: 770B

Hy4 preview is a flagship Mixture-of-Experts model designed for advanced reasoning, coding, agentic workflows, and efficient long-context processing. It features 770B total parameters with 49B activated per token, using 78 layers with 6,144 hidden size and 64 attention heads. The architecture combines Gated DeepSeek Sparse Attention with IndexCache, 256 routed experts and 1 shared expert, activating 8 routed experts per token, while iHC expands residual information flow across 4 streams. It includes a native Multi-Token Prediction layer for speculative decoding and supports up to a 1M-token context window.

Text GenerationInstruction Following+7
Runs onB200+

Z.ai GLM 5.3

Z.ai

Params: 320B – 753B

GLM 5.3 is an advanced Mixture-of-Experts model designed for long-horizon reasoning and large-scale coding, delivering substantially improved performance over GLM-5.2 through post-training. Its architecture features 78 layers with hybrid linear and DeepSeek sparse attention, 64 attention heads, 64 KV heads, and 256 routed experts with 8 activated per token. The model incorporates IndexShare sparse attention to improve long-context efficiency and an enhanced multi-token prediction layer for faster speculative decoding. Supporting up to a 1M token context window, it is optimized for complex reasoning, extended coding workflows, and sustained long-context tasks.

Text GenerationInstruction Following+7
Runs onB200+

Alibaba Qwen 3.8

Alibaba

Params: 27B – 2.4T

Qwen3.8 Flash Next is a 125B parameter hybrid model with 6B activated parameters, along with 51B n-gram embedding and 4B MTP parameters, designed for efficient long-context and agentic workloads. Its architecture combines Gated DeltaNet with Qwen Sparse Attention, 512 experts with 10 routed and 1 shared expert, Gated Residuals, and n-gram embeddings for efficient scaling and inference. It supports 262K-token context, extensible to 1M tokens, with native vision and video understanding. The model targets efficient deployment while reducing long-context latency and memory demands.

Text GenerationInstruction Following+8
Runs onB200+

IBM Granite 4.2

IBM

Params: 3B – 30B

Granite 4.2 30B is the flagship reasoning model in the Granite 4.2 family, designed for advanced coding, mathematics, tool calling, agentic workflows, and multilingual dialogue. It features 30B parameters with a 64-layer dense transformer using 4,096 hidden size, 32 attention heads, and 8 KV heads. The architecture employs Grouped Query Attention, RoPE with a 50M theta, and a 32,768-dimensional SwiGLU feed-forward network, with native reasoning and flexible thinking modes for balancing quality and latency. Supporting 128K native context with extension to 512K, it is optimized for complex long-context reasoning and enterprise applications.

Text GenerationInstruction Following+7
Runs onB200+

DeepReinforce.AI Ornith 1.5

DeepReinforce.AI

Params: 9B – 397B

Ornith 1.5 35B A3B is the mid-size Mixture-of-Experts member of the Ornith-1.5 family, designed for coding and agentic workloads with only ~3B parameters activated per token. Its end-to-end self-improvement approach jointly optimizes task generation, scaffold construction, and solution rollouts through reinforcement learning. The Qwen3.5-based architecture features 40 layers with hybrid linear and full attention, 16 attention heads, 2 KV heads, and 256 experts with 8 activated per token. The model supports a 262,144 token context window.

Text GenerationInstruction Following+7
Runs onB200+

Liquid AI LFM2.5 VL

Liquid AI

Params: 450M – 3B

LFM2.5 VL 3B is a compact multimodal language model designed for efficient on-device vision-language tasks, combining the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It uses a 30-layer hybrid architecture with 2,048 hidden size, 32 attention heads, and 8 key-value heads, alternating short convolution and full-attention layers for efficient sequence processing. Supporting a 32K-token context window, it uses native-resolution image processing with 512×512 patches and thumbnail processing. The model is optimized for image grounding, object detection, full-page OCR, layout understanding, multilingual applications, and efficient on-device multimodal inference.

Text GenerationInstruction Following+7
Runs onB200+

NVIDIA Nemotron v3

NVIDIA

Params: 30B – 550B

NVIDIA Nemotron 3.5 Lightning 30B A3B is a 30B-parameter hybrid Mixture-of-Experts model with 3B active parameters, combining Mamba-2, MoE, and Transformer attention layers. It features 52 layers, 128 routed experts with 6 activated per token, 32 attention heads, and a 2-head KV configuration, supporting up to 1M-token context, with multi-token prediction and reinforcement learning for reasoning, coding, tool use, and agentic workflows.

Text GenerationInstruction Following+9
Runs onB200+

Meta Muse Glimmer

Meta

Params: 30B

Muse Glimmer 30B is a multimodal dense causal language model designed for autonomous agentic tasks, combining multi-step reasoning, reliable tool use, failure recovery, and visual understanding. It has 30B parameters, including a dedicated 1.8B-parameter perception encoder, with a 52-layer language model using a 6,656 hidden size and 32 attention heads. The language model uses a repeating 3:1 local-to-global attention pattern with a 2,048-token sliding window and GQA with 32 query and 2 KV heads. It supports text and image inputs with a 131K+ context window and is optimized for local deployment and long-horizon agent workflows.

Text GenerationInstruction Following+8
Runs onB200+

Thinking Machines Inkling

Thinking Machines

Params: 276B – 975B

Inkling Small is a multimodal Mixture-of-Experts model with 276B total parameters and 12B active parameters, designed for conversational, coding, retrieval, and agentic applications. It features a 42-layer decoder-only architecture with 32 attention heads, 256 routed experts activating 6 per token, and 2 shared experts. Supporting text, image, and audio inputs, it uses hybrid local and global attention with a 1M-token context window and a 512-token sliding window, while supporting BF16 and NVFP4 numerics.

Text GenerationInstruction Following+8
Runs onB200+

Moonshot AI Kimi K3

Moonshot AI

Params: 2.8T

Kimi K3 is a frontier-scale native multimodal Mixture-of-Experts model designed for long-horizon reasoning, agentic coding, and large-scale knowledge workflows. It is the world's first open 3T-class model, featuring 2.8T total parameters with 104B activated, built on a 93-layer architecture with Kimi Delta Attention (KDA), Gated MLA, and Stable LatentMoE. The model uses a 7,168 hidden size, 96 attention heads, and 896 routed experts, activating 16 experts per token alongside 2 shared experts. Powered by MoonViT-V2 for native vision understanding and supporting a 1M-token context window, Kimi K3 is optimized for multimodal reasoning, long-context coding, advanced agentic workflows, and frontier-scale research applications

Text GenerationInstruction Following+8
Runs onMI355X+

Mindlab Research Macaron V1

Mindlab Research

Params: 35B – 753B

Macaron V1 Venti is a 753B parameter Mixture of LoRA (MoL) flagship model built on a 744B GLM-5.2 base with four 1B-parameter specialists for chat, personal-agent tasks, coding, and Generative UI. It uses a 78-layer architecture with a 6,144 hidden size and 64 attention heads, incorporating sparse MoE layers with 256 routed experts and 1 shared expert, activating 8 routed experts per token. The model integrates IndexShare sparse attention and multi-token prediction, while supporting a 1M-token context window for long-horizon agentic workflows, coding, tool use, and Generative UI.

Text GenerationInstruction Following+7
Runs onB200+

Poolside Laguna S 2.1

Poolside

Params: 118B

Laguna S 2.1 is a Mixture-of-Experts language model designed for agentic coding, reasoning, and long-horizon workflows. It features 118B total parameters with approximately 8B activated per token, using a 48-layer architecture with 3,072 hidden size and 48 attention heads. The model activates 10 experts per token across 256 routed experts alongside a shared expert, combining hybrid sliding window and full attention with grouped-query attention and per-head gating. Supporting up to a 1M token context window, it is optimized for large-scale coding, tool use, and long-context reasoning.

Text GenerationInstruction Following+7
Runs onB200+

Tencent Hy3

Tencent

Params: 295B

Hy3 is a Mixture-of-Experts (MoE) language model designed for advanced reasoning, coding, agentic workflows, and long-context processing. It features 295B total parameters with 21B activated, using an 80-layer transformer with 4,096 hidden size and 64 attention heads. The architecture includes 192 routed experts with top-8 routing alongside a shared expert, plus a Multi-Token Prediction layer for faster generation. Supporting up to a 256K context window, it is optimized for reliable tool use, complex multi-turn interactions, and high-throughput production inference.

Text GenerationInstruction Following+7
Runs onB200+

DeepReinforce.AI Ornith 1.0

DeepReinforce.AI

Params: 9B – 397B

Ornith 1.0 397B is a 397B parameter sparse Mixture-of-Experts model designed for high-performance agentic coding and software-engineering workflows. Built on Qwen3.5, it uses a 60-layer hybrid architecture combining linear and full attention, with a 4,096 hidden size, 32 attention heads, 2 KV heads, and 512 experts activating 10 per token. Approximately 17B parameters are active per token, with a 262K-token context window and multimodal support. Its self-improving RL framework jointly optimizes solution rollouts and scaffolds, enabling stronger search trajectories and higher-quality coding solutions.

Text GenerationInstruction Following+7
Runs onB200+

Alibaba Qwen AgentWorld

Alibaba

Params: 35B

Qwen AgentWorld 35B A3B is a native language world model designed to simulate agent environments across MCP, Search, Terminal, SWE, Android, Web, and OS interaction domains. It features 35B total parameters with 3B activated, using a 40-layer hybrid architecture with 2,048 hidden size and 16 attention heads. The model combines Gated DeltaNet and Gated Attention layers in a 3:1 pattern with 256 routed experts, activating 8 experts per token alongside a shared expert. Trained through CPT, SFT, and RL, it supports native world modeling and a 262K-token context window.

Text GenerationInstruction Following+7
Runs onB200+

Vultr VultronRetriever

Vultr

Params: 0.8B – 8B

VultronRetrieverPrime Qwen3.5 8B is a multimodal late-interaction retrieval model designed for visual document search and multilingual RAG across PDFs, scans, slides, and reports. It features 8B parameters, using a 32-layer hybrid GatedDeltaNet and full-attention architecture with 4,096 hidden size, 16 attention heads, and 320-dimensional multi-vector embeddings with MaxSim scoring. Supporting up to 262K context and 1,792 visual tokens, it delivers state-of-the-art retrieval accuracy while maintaining a compact index and efficient large-scale serving.

Semantic Search & RetrievalLate-Interaction Retrieval+5
Runs onB200+

Z.ai GLM 5.2

Z.ai

Params: 753B

GLM 5.2 is an advanced Mixture-of-Experts language model designed for long-horizon reasoning, large-scale coding, and million-token context processing. It features a 78-layer architecture with 6,144 hidden size and 64 attention heads, activating 8 experts per token across 256 routed experts and a shared expert. The model incorporates IndexShare sparse attention to improve long-context efficiency and an enhanced multi-token prediction layer for faster speculative decoding. Supporting up to a 1M token context window, it is optimized for complex reasoning, extended coding workflows, and sustained long-context tasks.

Text GenerationInstruction Following+7
Runs onB200+