AMD Instinct MI325X
8 GPUs per node
256 GB HBM3E
per GPU 2 TB Total
6.0 TB/s
per GPU throughput
CDNA 3
gfx942
6.4.2
Verified version
0.14.1
Verified version
This cookbook provides tested, working configurations for deploying LLMs on AMD hardware:
All configurations in this cookbook have been verified on:
8 GPUs per node
per GPU 2 TB Total
per GPU throughput
gfx942
Verified version
Verified version
Get a model running in under 5 minutes:
Test the endpoint:
From our comprehensive testing on MI325X:
Multi-run means (n=5). Peak throughput measured at optimal concurrency per model.