Cold start latency is one of the most underestimated costs in heterogeneous accelerator deployments. When a serving node receives its first inference request after model load, it performs a series of initialization operations that have nothing to do with the inference itself: CUDA/ROCm context initialization, kernel compilation, KV-cache preallocation, and attention pattern warmup. On H100 with CUDA, these operations are fast because the kernel cache is warm and the toolchain has been optimized for this path for years. On MI300X with ROCm, they take longer for reasons that are structural, not accidental.
We measured cold start latency for Llama 3 8B and 70B across H100 SXM and MI300X nodes under controlled conditions. The results: MI300X cold start for a fresh process takes 40 to 60 percent longer than H100 for the same model at comparable precision, depending on model size and whether JIT kernel compilation is triggered. This post explains the mechanism, the measurement methodology, and the mitigation strategies that actually reduce the impact in practice.
Why MI300X Cold Start Takes Longer
The primary driver is HIP kernel compilation. ROCm uses Just-In-Time compilation for many attention and matrix multiplication kernels during the first forward pass. When the kernel cache is empty (fresh container, first startup, or cache invalidated by a ROCm version update), the first forward pass must compile kernels before it can execute. On MI300X with Llama 3 70B BF16, this JIT compilation adds approximately 45 to 90 seconds to the initialization path in our measurements.
CUDA on H100 has a larger pre-compiled kernel library for serving workloads. Flash Attention 2, the attention kernels in vLLM, and the GEMM dispatch routines all have pre-built binaries for H100 that do not require JIT compilation. The cold start on H100 is dominated by model weight loading into HBM, not kernel compilation.
The second factor is HIP context initialization overhead. The ROCm runtime initialization for a fresh process adds a fixed cost that is higher than CUDA context init on equivalent hardware. We measured this at approximately 8 to 12 seconds for MI300X versus 3 to 5 seconds for H100, measured from process start to first successful GPU operation. For processes that restart infrequently, this is noise. For serving pods with aggressive restart policies or auto-scaling that spins up new pods on demand, it adds up.
Cold Start Latency Numbers by Configuration
Our measurements used vLLM 0.4.x, with warm (pre-existing kernel cache) and cold (fresh container, empty cache) configurations. Model loading latency (weight transfer to HBM) was measured separately from kernel init overhead.
For Llama 3 8B BF16:
- H100 cold start (process cold, weight loading): approximately 18 to 22 seconds
- MI300X cold start with warm ROCm kernel cache: approximately 24 to 30 seconds
- MI300X cold start with cold ROCm kernel cache (JIT compilation): approximately 65 to 95 seconds
For Llama 3 70B BF16, single node:
- H100 TP=2 cold start: approximately 55 to 70 seconds (dominated by 140GB weight loading across 2 nodes)
- MI300X single node cold start with warm cache: approximately 72 to 88 seconds (192GB weight loading)
- MI300X single node cold start with cold cache: approximately 140 to 180 seconds
These are process-level cold start times, not request-level latency. The first inference request after cold start will see latency that includes both the initialization overhead and the actual inference time. Subsequent requests on a warm process see only inference latency.
ZML Warmup Scheduling: How We Reduce the Impact
ZML implements a warmup scheduler that runs synthetic inference passes against a newly registered node before routing live traffic to it. The warmup pass is a configurable sequence of forward passes at different batch sizes and sequence lengths, designed to trigger the JIT kernel compilations that would otherwise be triggered by the first live request.
The warmup policy configuration includes a warmup_passes field that specifies the sequence of synthetic inputs to run, and a ready_threshold that defines the criteria for marking a node ready for traffic:
warmup:
enabled: true
warmup_passes:
- batch_size: 1
seq_len: 128
- batch_size: 8
seq_len: 256
- batch_size: 32
seq_len: 512
ready_threshold:
max_init_latency_ms: 2000
min_consecutive_successes: 3
With warmup enabled, a new MI300X node is excluded from the live routing pool until the warmup sequence completes successfully. This means the first live request to the node arrives after JIT compilation has run, and the first-request latency reflects only inference latency, not compilation overhead.
The cost is extended time-to-ready for new nodes: instead of a node becoming available 30 seconds after process start, it becomes available after the warmup sequence completes, which may take 90 to 120 seconds for a cold MI300X process with full warmup. For long-running serving pods that restart rarely, this overhead is paid once per node lifetime. For auto-scaling scenarios where new pods are created to handle traffic spikes, the warmup time is the latency before the new capacity is usable, which needs to factor into scaling lead time.
Practical Mitigation Strategies
Beyond ZML's built-in warmup scheduling, there are three other mitigations that materially reduce cold start impact.
ROCm kernel cache persistence: the JIT-compiled kernel cache can be persisted to a mounted volume and pre-loaded on container startup. When the kernel cache is warm (mounted from a pre-compiled state), MI300X cold start drops to the weight-loading-only path, which is much closer to H100 in relative terms. This requires building the cache from a representative warmup run and managing it as an artifact, but it is the highest-impact single change for teams running MI300X pods with significant restart frequency.
Keep-warm pods: maintaining at least one warm, idle MI300X node per accelerator pool rather than scaling to zero prevents the cold kernel cache scenario on scale-up events. This is a cost tradeoff: you pay for an idle pod to avoid cold start penalty on demand spikes. For workloads with predictable traffic patterns, this is often cost-justified.
Startup probe configuration: Kubernetes readiness probes should be configured with initial delay and period settings that account for actual cold start time. A probe interval of 10 seconds with an initial delay of 30 seconds will correctly avoid marking an MI300X pod ready before initialization completes. The default probe settings in many serving deployments are calibrated for H100 startup times and will incorrectly mark MI300X pods as ready during kernel compilation, routing traffic to a node that will respond slowly or time out on the first request.
Cold Start Is a Solvable Engineering Problem
The 40 to 60 percent cold start gap between MI300X and H100 is real and consequential for auto-scaling workloads. It is not, however, a reason to avoid MI300X for serving deployments. The gap is addressable with warmup scheduling, kernel cache persistence, and appropriate health probe configuration. Teams that implement these mitigations are running MI300X in production without users experiencing the initialization overhead. Teams that do not implement them will encounter it as a user-facing latency anomaly and spend time diagnosing what is actually a known and predictable behavior.
ZML routes around cold nodes automatically. No SLA impact during kernel compilation.
ZML health probes track warm vs cold state per node and route new requests only to warm instances.