Minimum and recommended specifications for running vLLM on AMD Instinct GPUs.
Both GPUs share the same architecture and use identical vLLM configurations.
LLM inference performance scales with memory bandwidth, not compute. MI325X's 6.0 TB/s bandwidth directly translates to higher token throughput, especially for large batch sizes where memory access patterns dominate.
Specifications from AMD MI325X product page.
Memory Advantage
This cookbook was tested with the following versions:
For multi-GPU systems, NUMA balancing must be disabled:
If enabled, disable it:
0 Comments
Be the first to comment and share your perspective with the community.