Over 80,000,000 Cloud Servers Launched

Over 80,000,000 Cloud Servers Launched

Cloud ComputeCloud GPUBare MetalFile SystemObject StorageBlock StorageManaged DatabasesCDNServerlessKubernetesContainer RegistryDirect ConnectLoad Balancers
RegionsAdvanced NetworkControl PanelOperating SystemsUpload ISO
Industry CloudOne-Click DeploymentUse Cases
Browse AppsBecome a Vendor
FAQDevelopers / APIsVultr DocsServer StatusBug BountyPromotionsSolution PartnersStart-Up Programs
Our TeamNewsBrand AssetsReferral ProgramCreator ProgramCareersSLALegalVultr Trust CenterContactYour Privacy ChoicesSubprocessorsAccessibility
Contact SalesSign Up

© Vultr 2026 | VULTR is a registered trademark of The Constant Company, LLC.

Terms of ServiceAUPDMCAPrivacy PolicyCookie Policy
  • Pricing
DashboardContact Sales
  • Blogs
  • Discover
  • Docs
  • Community
⌘K
Inference Cookbook
  1. Docs
  2. Inference Cookbook

Inference Cookbook

Comprehensive inference cookbook for running large language models on NVIDIA HGX B200 and AMD Instinct GPUs using vLLM.

Inference Cookbook for CUDA

See All
GPU

NVIDIA HGX B200

8 GPUs per node

VRAM

179 GB HBM3e

per GPU 1.44 TB Total

Interconnect

NVSwitch + NVLink 5.0

1.8 TB/s bidirectional

Getting Started

Get started with NVIDIA HGX B200 GPUs, including hardware overview, environment setup, and first model deployment.

Model Guides

Deploy leading AI models including Nemotron, DeepSeek, GLM, and MiniMax on NVIDIA HGX B200 GPUs with optimized inference configurations.

Dynamo

Explore NVIDIA Dynamo’s architecture for disaggregated LLM inference, including routing, KV cache tiering, and optimized deployment with vLLM.

Optimization

Optimize LLM inference on NVIDIA HGX B200 GPUs with KV cache management, quantization, kernel tuning, and concurrency optimization.

Benchmarks

Benchmark methodology and performance results for LLM inference workloads on NVIDIA HGX B200 GPUs.

Troubleshooting

Common issues encountered when running vLLM on NVIDIA HGX B200 GPUs and their solutions.

Production Deployment

Guidelines for deploying vLLM on NVIDIA HGX B200 instances in production.

Inference Cookbook for ROCm

See All
GPU

AMD Instinct MI325X

8 GPUs per node

VRAM

256 GB HBM3E

per GPU 2 TB Total

Architecture

CDNA 3

gfx942

Getting Started

Get started running vLLM on AMD Instinct GPUs with hardware requirements, environment setup, and your first model deployment.

Model Guides

Deploy large AI models including DeepSeek V3.2, Llama 3.1, Qwen3-VL, and Kimi-K2.5 on AMD Instinct GPUs.

Optimization

Improve LLM inference performance on AMD Instinct GPUs with FP8 quantization, KV cache optimization, and concurrency tuning.

Benchmarks

Detailed benchmarking of DeepSeek, Llama, Qwen3-VL, and Kimi models on AMD Instinct MI325X GPUs with stress and validation testing.

Troubleshooting

Common issues and verified solutions for vLLM on AMD Instinct GPUs.

Production Deployment

Deploy vLLM with health checks, monitoring, and resilience on AMD Instinct GPUs.

Popular Articles

Troubleshooting

11 March 2026

Production Deployment

11 March 2026

Troubleshooting

17 March 2026

Production Deployment

11 March 2026

Tech Talk: Running a Self-Improving AI Agent with Hermes on Vultr

From The Blog

Vultr at All-In Summit 2026: Connecting the Future of Enterprise AI Infrastructure

October 5, 2026

Creators(x) Spotlight: Mirdul Swarup – Building in the Open

October 1, 2026

Stealthium Joins the Vultr Cloud Alliance to Bring GPU Observability and Runtime Security to AI Infrastructure

September 30, 2026

View all posts