Deploy Meta's Llama-3.1-405B-Instruct on AMD Instinct GPUs.
Llama-3.1-405B requires accepting Meta's license agreement on HuggingFace. Visit the model page to request access.
Multi-run means (n=5).
See Llama-3.1-405B Stress Testing for detailed results.
For 128K context (requires more memory):
For tighter memory constraints:
Reduce context length:
Enable FP8 quantization for better performance:
0 Comments
Be the first to comment and share your perspective with the community.