OCI GPU Quick Start: AMD MI300X Instinct GPU

This document provides hardware specifications, supported OS images, onboarding verification, sample benchmarks, and best-practices for OCI deployments using the AMD MI300X GPU bare-metal shape.

BM.GPU.MI300X.8 is a scale-up and scale-out OCI GPU shape built around eight AMD Instinct MI300X GPUs with XGMI connectivity, dual Intel Xeon Platinum 8480+ processors, and RoCE-capable networking for AI/ML and HPC workloads.

At a Glance

  • Shape: BM.GPU.MI300X.8
  • GPU configuration: 8 x AMD Instinct MI300X 192 GB
  • Recommended OS baseline: Oracle Linux 9+ or Ubuntu Linux 22.04+
  • Recommended software baseline: ROCm 7.2+, RCCL 2.26.6, OFED 28.40.1202+
  • Primary verification command: amd-smi
  • Operational profile: single-node and multi-node AI/ML and HPC with XGMI topology

Table of Contents

Hardware Specifications

Shape Name GPU Model GPUs/Node GPU Memory (GB/GPU) GPU Memory Total CPU # of CPUs System Memory Local Storage Host NIC RDMA (RoCE) NICs
BM.GPU.MI300X.8 AMD Instinct MI300X 8 192 1.5 TB 2 x Intel Xeon Platinum 8480+ 112 Cores 2 TB DDR5 8 x 3.5 TB NVMe = 28 TB 100 Gb/s 8 x 1 x 400 Gb/s = 3.2 Tb/s

See the OCI Compute Shapes Docs for up-to-date details.

Recommended Operating Systems

  • Oracle Linux 9+
  • Ubuntu Linux 22.04+
  • ROCm 7.2+
  • RCCL 2.26.6
  • OFED 28.40.1202+
  • Oracle Cloud Agent 1.52.0+
  • HPC-X 2.23+

Custom OS image Creation with Packer

To build your images using packer clone the OCI HPC Images repo and run the commands found there OCI HPC Images GitHub Repo.

Provided Images

It is recommended to use Oracle-provided or organizationally-approved images for security and supportability. See the OS images table for region-specific OCIDs and version information.

OS Version Image Packer Build Details OCI Platform Image Link Driver Versions
OCI GPU AI Image with Oracle Linux 9 Oracle-Linux-9.6-RHCK-DOCA-OFED-3.2.1-AMD-ROCM-643 PAR Link ROCm 6.4.3, DOCA OFED 3.2.1, OCA 1.57.0
OCI GPU AI Image with Oracle Linux 9 Oracle-Linux-9.7-RHCK-DOCA-OFED-3.2.1-AMD-ROCM-72 PAR Link ROCm 7.2, DOCA OFED 3.2.1, OCA 1.57.0
OCI GPU AI Image with Ubuntu Linux 22.04 Canonical-Ubuntu-22.04-DOCA-OFED-3.2.1-AMD-ROCM-643 PAR Link ROCm 6.4.3, DOCA OFED 3.2.1, OCA 1.57.0
OCI GPU AI Image with Ubuntu Linux 22.04 Canonical-Ubuntu-22.04-DOCA-OFED-3.2.1-AMD-ROCM-72 PAR Link ROCm 7.2, DOCA OFED 3.2.1, OCA 1.57.0
OCI GPU AI Image with Ubuntu Linux 24.04 Canonical-Ubuntu-24.04-6.8-DOCA-OFED-3.2.1-AMD-ROCM-643 PAR Link ROCm 6.4.3, DOCA OFED 3.2.1, Kernel 6.8, OCA 1.57.0
OCI GPU AI Image with Ubuntu Linux 24.04 Canonical-Ubuntu-24.04-6.8-DOCA-OFED-3.2.1-AMD-ROCM-72 PAR Link ROCm 7.2, DOCA OFED 3.2.1, Kernel 6.8, OCA 1.57.0
OCI GPU AI Image with Ubuntu Linux 24.04 Canonical-Ubuntu-24.04-6.14-DOCA-OFED-3.2.1-AMD-ROCM-72 PAR Link ROCm 7.2, DOCA OFED 3.2.1, Kernel 6.14, OCA 1.57.0

Hello World Verification

Use amd-smi or rocm-smi to confirm that all eight GPUs are visible and idle before launching a workload:

amd-smi
rocm-smi

amd-smi topology should show a fully connected XGMI topology across the eight devices.

Performance Benchmarks

For benchmark setup guidance, see the AMD RCCL documentation and the AMD system validation guidance.

RCCL & Model Inference Performance

The current MI300X source material emphasizes RCCL, RVS, and system-validation workflows more than a normalized model-inference table. Collective communication results below are sourced from the current handover document.

All Reduce - Single Node

  • Large-message reference point: 188.65 GB/s algorithm bandwidth and 330.13 GB/s bus bandwidth at 17179869184 bytes

All Reduce - 2 Nodes

  • Large-message reference point: 200.11 GB/s algorithm bandwidth and 375.21 GB/s bus bandwidth at 17179869184 bytes

Model Inference Performance

No normalized MI300X model-inference table is currently published in the source material used for this workflow.

OKE GPU Getting Started

  1. Create an OKE cluster through the OCI Console or CLI.
  2. Add the GPU node pool using one of the approved MI300X images listed above.
  3. Install the AMD device plugin:
kubectl apply -f https://github.com/AMD/k8s-device-plugin/raw/main/AMD-device-plugin.yml
  1. Verify that the GPUs are visible from a test pod before deploying production workloads.

Troubleshooting

Here you can find suggested troubleshooting methods.

Confirm Hardware & Drivers

amd-smi
amd-smi topology
numactl --hardware

The MI300X source material expects all eight GPUs to be visible and amd-smi topology to report a symmetric XGMI topology.

rdma link

The front-end network is expected on mlx5_1 (eth0) with RoCE RDMA interfaces exposed as rdma* devices and links reporting ACTIVE and LINK_UP.

RVS

/opt/rocm/bin/rvs

RVS is included with the ROCm stack and is the preferred system-validation workflow called out by the source material for MI300X.

Further Reading & Support

AMD Technical Documents

  1. AMD ROCm compatibility matrix
  2. AMD ROCm examples
  3. AMD ROCm Validation Suite
  4. AMD system validation guidance

OCI AMD Technical Docs & Supported Solutions

  1. OCI GPU Scanner - Cluster Health Management Solution
  2. OCI HPC SLURM Deployed Terraform Stack
  3. OCI HPC OKE Deployment Stack