
AMD AI Workbench is a low-code AI development platform that simplifies fine-tuning, inference, and other AI jobs on AMD Instinct™ GPUs. It includes a comprehensive model catalog, an AIM Catalog for deploying AMD Inference Microservices, and integrations with MLflow, TensorBoard, and Kubeflow. The Vultr Marketplace provides a pre-configured AMD AI Workbench application that deploys directly into a Vultr Kubernetes Engine (VKE) cluster running on AMD Instinct™ MI3xx GPUs, enabling rapid platform setup for AI teams.
This guide explains deploying and using Vultr's AMD AI Workbench VKE Application on Vultr Kubernetes Engine. It covers adding a bare-metal AMD Instinct™ GPU node pool to an existing VKE cluster, deploying the AMD AI Workbench application through the VKE Applications interface, and accessing the AMD Resource Manager and AMD AI Workbench UIs to manage AI workloads.
Before you begin, ensure you:
kubectl configured on your workstation to access the cluster by following the Vultr Kubernetes Engine connection guide.Add a bare-metal GPU node pool with AMD Instinct™ MI3xx GPUs to the cluster. The AMD AI Workbench application schedules its inference and fine-tuning workloads on these GPU nodes.
From the Kubernetes section, click the cluster name to open its details page.
Click the Node Pools tab.
Click Add Node Pool.
Configure the GPU node pool.
Click Create Node Pool to provision the GPU nodes.
Bare-metal GPU provisioning takes longer than virtual nodes. Allow 15 to 30 minutes for the nodes to register with the cluster.
List the registered nodes to verify that the GPU nodes joined the cluster.
The output displays both the CPU and GPU node pools with the GPU nodes showing the Ready status.
After the GPU nodes are Ready, deploy the AMD AI Workbench VKE Application from the VKE Applications interface. The application installs the AMD Enterprise AI Reference Stack components, including AMD AI Workbench, AMD Resource Manager, Kaiwo, Cluster Forge, and AIMs.
From the cluster details page, click the Applications tab.
Click Deploy Application.
Select AMD AI Workbench from the available VKE Applications.
Set the Timeout to 60 minutes to allow sufficient time for all platform components to install.
Click Deploy Now to start the application deployment.
Application deployment runs two background jobs in the amd-ai-system namespace. Monitor both jobs as described below.
The VKE Application launches two Kubernetes jobs that run in sequence: a GPU node configuration job followed by the main bootstrap job. Follow each job's logs to confirm the deployment is progressing and to surface errors early.
Follow the GPU node configuration job logs. This job runs first and typically completes in 1 to 2 minutes.
After the GPU node configuration job completes, follow the bootstrap job. This is the main installation and typically takes 30 to 35 minutes.
The bootstrap logs display the auto-generated domain and load balancer IP in a block similar to the one below.
The VKE Application uses nip.io to generate a wildcard-resolvable domain from the load balancer IP. No DNS configuration is required.
Wait for the completion marker in the bootstrap logs.
Save the domain name for use in the next section.
After the bootstrap job completes, the platform exposes several UIs and APIs at predictable subdomains under the auto-generated nip.io domain. The following table lists the available service endpoints. Replace 192-0-2-10 in the URLs with the dashed form of your load balancer IP from the bootstrap logs.
Retrieve the initial devuser password.
Open the AMD AI Workbench UI or the AMD Resource Manager UI in your browser and sign in with the devuser credentials.
https://aiwbui.amd-ai-suite.192-0-2-10.nip.iohttps://airmui.amd-ai-suite.192-0-2-10.nip.iodevuser@amd-ai-suite.192-0-2-10.nip.ioRetrieve the Keycloak admin password to access the admin console.
Open the Keycloak admin console in your browser and sign in.
https://kc.amd-ai-suite.192-0-2-10.nip.iosilogen-adminAMD AI Workbench supports a range of AI workloads on Vultr's AMD Instinct™ GPU infrastructure.
In this guide, you deployed Vultr's AMD AI Workbench VKE Application on a Vultr Kubernetes Engine cluster. You added a bare-metal AMD Instinct™ GPU node pool, installed the application through the VKE Applications interface, and accessed the AMD AI Workbench, AMD Resource Manager, and Keycloak admin UIs. With AMD AI Workbench's model catalog, AIM Catalog, and integrated MLOps tooling on Vultr's AMD Instinct™ GPU infrastructure, you can run fine-tuning, inference, and multi-tenant AI workloads on a unified platform. For more information, visit the official AMD AI Workbench documentation.
0 Comments
Be the first to comment and share your perspective with the community.