Sign In Get Started
Back to Blog

Running LLM Inference on AMD MI300X: What Infrastructure Teams Need to Know

AMD MI300X offers compelling cost-per-token for large batch workloads, but the path from NVIDIA to AMD is rarely smooth. We document what actually changed when we moved Llama 3 70B workloads across to MI300X hardware.

AMD MI300X accelerator hardware in a production server rack environment

AMD MI300X has been live in major cloud providers since late 2023, and the hardware specs make a real case for it: 192GB of unified HBM3 memory, peak BF16 throughput that beats H100 on paper, and spot pricing that frequently sits 20 to 40 percent below equivalent H100 nodes. When we started shifting production Llama 3 70B workloads to MI300X nodes in our test cluster, we expected the migration to be mostly mechanical. It was not.

This post covers what we actually found, layer by layer, from the ROCm software stack to attention kernels to precision format behavior. We are writing this for infrastructure engineers evaluating the NVIDIA-to-AMD move for real production workloads, not synthetic benchmarks.

The ROCm Software Stack: What You Need to Audit First

The most consequential difference between NVIDIA and AMD production inference is the software underneath. CUDA toolchains have years of production hardening and a large ecosystem of pre-compiled serving kernels. ROCm is a legitimate production runtime but has a shorter track record in serving workloads specifically, and the HIP translation layer introduces failure modes that only surface under load.

When we configured vLLM with ROCm support (0.4.x branch with the ROCm build path), we encountered two categories of problems. First, not all custom CUDA kernels have HIP equivalents that have gone through the same tuning cycles. Flash attention on ROCm ships with different performance characteristics than on CUDA, and the gap is not always visible in documentation. Second, driver and toolkit version pinning matters more on ROCm: a minor toolkit mismatch between the host driver and the container runtime caused silent performance degradation on one of our node configurations. We caught it through latency tracking, not through any explicit error.

The ROCm ecosystem has improved significantly through 2024 and 2025, and the core serving path for large models is stable today. The point is that any migration requires a deliberate software inventory step before you change a single inference call. Pin your ROCm version. Validate your kernel coverage. Run a load test before switching production traffic.

The 192GB Memory Advantage Has Conditions

The MI300X ships with 192GB of unified HBM3, versus 80GB on H100 SXM5. For Llama 3 70B in BF16, that means a single MI300X node can hold the full model weights without tensor parallelism, while H100 typically requires two nodes for the same configuration. For large-model workloads at smaller batch sizes, where inter-GPU communication overhead is a real cost, this is a genuine advantage.

Two conditions limit it in practice. The unified memory architecture on MI300X means peak memory bandwidth is shared across all compute units. Under sustained inference load with large batch sizes, we measured effective memory bandwidth utilization meaningfully below the hardware spec, particularly for attention operations accessing memory in irregular patterns. The second condition is KV-cache management: when the entire model fits in one accelerator's memory, you gain the GPU communication savings but lose some of the KV-cache paging strategies that vLLM has tuned for multi-GPU configurations. Whether that tradeoff favors MI300X depends on your typical sequence length and batch composition.

Precision Format Behavior: BF16, FP8, and FP16

We tested Llama 3 70B at three precision formats: BF16, FP8-E4M3, and FP16. The results were not uniform across hardware families.

BF16 performance on MI300X matched our expectations from published hardware specs. FP8 inference was more variable. The MI300X supports FP8 natively, but as of our test period, FP8 kernel implementations in the major serving frameworks were still being tuned for the ROCm path. We saw cases where FP8 on MI300X was slower than BF16 on the same hardware, which is the opposite of what you see on H100 SXM with mature FP8 kernels. FP16 was effectively equivalent to BF16 for our workloads, as expected.

The relevant caution: if your primary motivation for moving to MI300X is FP8 cost reduction, validate FP8 kernel maturity for your specific framework and ROCm version before committing. Hardware support for FP8 and production-quality kernel coverage for FP8 are two different things.

Throughput vs. Latency: The Split That Determines Where MI300X Belongs

For Llama 3 70B at batch sizes above 32, we found MI300X throughput competitive with H100 SXM. At batch sizes above 64, MI300X showed higher tokens-per-second for prefill-heavy workloads, consistent with the larger memory bandwidth and unified pool. This is the sweet spot where the hardware wins.

For latency-critical paths at small batch sizes (below 16 sequences), H100 SXM maintained a consistent edge. CUDA kernel optimization for attention at small batch sizes reflects years of inference-specific tuning that ROCm has not fully replicated. For teams running mixed traffic, some latency-sensitive and some throughput-oriented, this points toward routing strategy rather than wholesale hardware replacement.

We are not saying MI300X is wrong for latency-critical workloads. We are saying the cost-benefit calculation changes significantly at small batch sizes, and teams optimizing for p50 latency should benchmark their specific workload rather than assuming the headline FLOPS figure translates.

Four Things That Required Changes in Our Setup

The practical migration required changes in four areas. Containerization: we had been using NVIDIA NGC containers, and MI300X required building ROCm-based containers with pinned toolkit versions and explicit HIP library dependencies. Not difficult, but it adds a maintenance surface that did not exist before.

Health checks and warmup: MI300X nodes have longer initialization times than H100 for the first inference batch after a cold start. Our pod readiness probes were triggering premature traffic routing. We extended the warmup buffer, which resolved it, but the behavior needs to be understood before you hit it in production.

Batching parameters: the optimal max_num_seqs and gpu_memory_utilization settings for vLLM differed between hardware families. Settings tuned for H100 did not transfer directly to MI300X. This is not surprising but it means you need a tuning pass, not just a hardware swap.

Monitoring: MI300X exposes different hardware counters through ROCm SMI versus nvidia-smi. Our Prometheus exporters needed updates to pull equivalent metrics from AMD hardware. A routine change, but easy to skip until production visibility degrades.

The Practical Conclusion

MI300X has reached production viability for LLM inference, specifically for batch workloads above 32 sequences per step, large model serving where the 192GB memory pool eliminates a tensor parallelism tier, and cost-sensitive deployments where the spot price difference justifies migration work.

The migration is not a drop-in swap. It requires ROCm version management, kernel stack revalidation, batching retuning, and a warmup strategy update. For teams that want to hold both hardware families in their cluster without maintaining two separate inference pipeline configurations, the answer is abstracting the hardware assignment behind a routing layer. The serving framework stays the same; the accelerator is a runtime policy decision. That is the approach we built ZML around, and the MI300X migration is what made us understand why that abstraction has to happen at dispatch time, not at framework configuration time.

Route inference across H100 and MI300X without rewriting your pipeline

ZML connects your production models to H100, MI300X, Gaudi2, and more from a single API. Start free, no card required.

Read the quickstart Read the docs