Benchmark comparisons between H100 and MI300X have proliferated over the past year, and most of them share the same problem: they test at unrealistically large or small batch sizes, use precision formats that do not match production deployments, or benchmark frameworks that have not been tuned for the hardware under test. This post describes the setup we used for our internal measurements, the numbers we got, and what we think the results actually mean for production serving decisions.
We are publishing internal test cluster results, not a controlled third-party study. The hardware was H100 SXM5 80GB nodes and AMD MI300X 192GB nodes, both in our on-premise test cluster. The model was Llama 3 8B and Llama 3 70B. The serving framework was vLLM 0.4.x, with H100 using CUDA backend and MI300X using ROCm backend. We tested five batch sizes (1, 8, 32, 64, 128 concurrent sequences) and three precision formats (BF16, FP16, FP8).
Test Methodology
Prefill and decode throughput were measured separately. For prefill, we used a fixed prompt of 512 tokens and measured tokens-per-second for the forward pass across the batch. For decode, we measured tokens-per-second during autoregressive generation for 256 output tokens, averaging across the full batch at each batch size. Latency measurements are p50 and p95 for the full request (prefill plus decode), measured from request receipt to final token.
All FP8 results use E4M3 format, which is the dominant FP8 variant in production serving. H100 FP8 results used CUDA FP8 kernels with Flash Attention 2 FP8 support. MI300X FP8 results used ROCm FP8 support as available in the vLLM ROCm build at our test version. ROCm FP8 kernel coverage has been improving, and these numbers reflect our specific test date: exact results may differ with later ROCm and vLLM versions.
Llama 3 8B Results
For Llama 3 8B, H100 and MI300X performed similarly at small batch sizes, with H100 holding a modest latency advantage at batch size 1 and 8 for p95 decode. At batch sizes 32 and above, results converged. At batch size 128, MI300X prefill throughput exceeded H100 by approximately 12 percent in BF16, consistent with MI300X's higher memory bandwidth for this model size.
FP8 on H100 showed a 15 to 22 percent throughput improvement over BF16 at batch sizes 32 and above, which aligns with expected FP8 compute density gains on H100. FP8 on MI300X showed more modest gains: 5 to 10 percent at equivalent batch sizes, reflecting the current maturity gap in ROCm FP8 kernel optimization for this model family.
Key takeaway for 8B: for this model size, both hardware options are competitive in production. If your deployment is FP8-primary, H100 has a real advantage until ROCm FP8 kernels catch up. If you are running BF16 at high batch sizes, MI300X is a viable alternative at lower cost.
Llama 3 70B Results
The 70B results showed a more pronounced split between hardware families. For H100, running Llama 3 70B requires tensor parallelism across two 80GB nodes (TP=2). For MI300X, the 192GB pool fits the full model in BF16 on a single node, eliminating the GPU-to-GPU communication overhead of tensor parallelism.
At batch size 1 (single request), H100 TP=2 had lower p95 latency than single MI300X: roughly 18 percent faster for full request latency. This is the H100 sweet spot for interactive, latency-critical serving. At batch size 32, the numbers were close, within measurement noise for both throughput and latency. At batch sizes 64 and 128, single MI300X outperformed H100 TP=2 on prefill throughput by 8 to 14 percent, and matched on decode throughput.
Cost-normalized: when you price two H100 SXM5 nodes at spot rates versus one MI300X node at spot rates, the per-token cost for 70B at batch sizes above 32 favored MI300X by 20 to 28 percent in our test period. That gap is driven by spot pricing ratios, which fluctuate, but the direction has been consistent: MI300X spot has been below H100 spot for comparable workloads throughout 2025.
FP8 on 70B: The Cautionary Note
We tested FP8 E4M3 on Llama 3 70B on both hardware types. On H100, FP8 showed consistent 18 to 25 percent throughput improvement over BF16, across batch sizes and for both prefill and decode. On MI300X, FP8 performance was inconsistent: at batch size 32 we measured comparable performance to BF16, and at batch size 64 we measured performance below BF16 for decode by about 7 percent. This is not expected behavior for FP8 relative to BF16 and reflects kernel maturity issues on the ROCm path for this model size and batch range at our test ROCm version.
We want to be clear about what this means: the hardware itself supports FP8 at the silicon level. The gap is in serving framework kernel optimization for this specific combination of model size, batch shape, and ROCm version. This is an area of active improvement in the ROCm ecosystem. Teams planning MI300X FP8 deployments for 70B models should run their own validation against current ROCm and vLLM versions rather than relying on these or any other fixed-date benchmarks.
What These Numbers Actually Tell You
The clean summary from our internal data: H100 SXM wins for latency-critical 70B serving at batch size 1-8, and for any serving using mature FP8 kernels. MI300X is competitive or better for 70B batch processing at batch sizes 32 and above in BF16, and for 8B serving across most use cases, at meaningfully lower spot cost.
The less clean summary: your specific workload numbers will differ from ours based on traffic shape, context length distribution, ROCm version, and vLLM tuning. These numbers are a starting point for calibrating expectations, not a definitive production prediction. The only way to know your actual cost-performance tradeoff is to benchmark on your traffic with your serving configuration.
What the benchmark exercise confirmed for us: there is no single "better" hardware for LLM inference across all workloads. The right answer depends on batch size, model size, latency target, and cost priority, which means routing policy, not hardware monoculture, is how you handle the real production environment.
Let ZML route your Llama 3 workloads to whichever node wins the cost equation today
ZML tracks live cost-per-token for each accelerator in your pool and routes accordingly.