Between Q3 2023 and Q4 2025, H100 spot availability in us-east-1 dropped below 60 percent on at least six separate occasions, each lasting between four and twelve days. Teams that noticed were the ones watching their spot request queue depths. Teams that did not notice were the ones waking up to latency alerts and wondering why their inference service had gone from 200ms median to 4 seconds.
GPU supply volatility is structural, not cyclical. It is driven by a combination of fabrication lead times measured in months, data center build schedules that lag demand spikes by six to eighteen months, and allocation prioritization that changes when large cloud customers pre-buy capacity. Individual production teams have no visibility into any of these factors and no control over any of them. The question is whether you design around that reality or absorb the cost of ignoring it.
Why Supply Shocks Feel Like Software Bugs
When H100 spot availability drops, the first symptom is not a clear "no hardware available" error. It is an inference request queue that grows silently, increasing p99 latency while median latency stays stable for a short period, followed by a cascade when the queue depth exceeds the serving framework's configured limit. This looks like a software problem or a traffic anomaly, not a hardware supply event. Infrastructure engineers spend real hours debugging the wrong layer before tracing the root cause to spot capacity.
We have seen this pattern across multiple teams in our early-access program. The sequence is: latency spike, framework restart, latency spike recurs, escalation, infrastructure audit, eventually someone checks cloud console spot capacity and finds the request queue. The median time from symptom onset to correct diagnosis was around 3 hours in teams without explicit spot capacity alerting. With alerting in place, it was under 15 minutes. But most teams do not have spot capacity alerting set up because they have not been burned by this specific failure mode yet.
The Engineering Debt From Single-Accelerator Assumptions
A serving pipeline built for H100 only is not just unportable in the abstract. It accumulates specific forms of technical debt that become expensive when supply changes.
CUDA-specific code paths exist throughout the stack. Custom attention kernels compiled for H100. Quantization pipelines using TensorRT or equivalent tooling that generates H100-specific binaries. Monitoring instrumentation that calls nvidia-smi and has no AMD equivalent. Containerization that uses NVIDIA base images. None of these components are wrong; they are appropriate for the hardware. But each one is a migration cost that must be paid when the hardware changes.
The migration cost estimation problem is that these dependencies are not visible from the application layer. Teams estimate "a few days to swap hardware" because the inference call itself is a line or two of Python. The actual work is surfaced only when the migration starts: discovering the custom kernel, the TensorRT pipeline, the monitoring gap, the container rebuild. We tracked this across teams attempting emergency H100-to-MI300X migrations and found the mean engineering time was 2.8 weeks for a serving stack with moderate customization, versus the 3-5 days originally estimated.
Designing for Accelerator Portability from the Start
The right time to design for hardware portability is before the first supply shock, not after. The design principles are not complicated, but they require deliberate choices at the stack architecture stage.
Inference framework selection: serving frameworks that support multiple hardware backends (vLLM with ROCm support, TGI with Gaudi backends, Triton Inference Server with hardware abstraction) reduce the per-hardware migration surface compared to frameworks built for a single vendor. The tradeoff is occasionally paying a small performance cost on your primary hardware for kernels that were not tuned specifically for it. For most production serving workloads, that cost is less than the migration cost when supply changes.
Hardware abstraction in dispatch: the layer that routes an inference request to a specific GPU node should not have hardcoded hardware type assumptions. Dispatch policies (which hardware handles which request type, and in what fallback order) should be external configuration, not code. This is the architectural decision that makes supply-driven routing possible without code changes.
Precision format validation on all planned hardware: if you use FP8 or INT8 quantization, validate kernel support on your fallback hardware before you need it. FP8 support on AMD hardware has improved but is not at parity with CUDA FP8 kernels across all model families as of early 2026. Running a quantization validation before a supply shock beats discovering the gap during one.
What an Accelerator-Agnostic Serving Stack Looks Like in Practice
Consider a serving configuration for a Llama 3 70B deployment. The routing policy defines three tiers: H100 SXM nodes for requests with latency target under 300ms, MI300X nodes for batch requests with latency budget up to 2 seconds, and H100 PCIe on-demand as a fallback when spots on both are constrained. The model weights, quantization format, and serving framework configuration are identical across all three hardware types. The only hardware-specific configuration is the kernel priority hints passed to the serving framework at startup.
When H100 spot availability drops, the routing layer shifts traffic to MI300X without any code change or deployment. When MI300X has maintenance windows, traffic routes to on-demand PCIe nodes at a higher cost, but without downtime. The serving SLA is maintained. The engineering team is not paged.
We are not claiming this is trivial to set up. It requires validating your model's performance profile on each hardware type, tuning batching parameters per hardware, and maintaining a routing policy that reflects your cost and latency priorities. That is real work. But it is work you do once, and then the ongoing cost of supply fluctuations drops to near zero.
The Asymmetric Cost of Doing Nothing
The argument against investing in hardware portability is usually: "we have not had a supply problem yet." This is a base rate argument that ignores severity and asymmetry. The probability of a major H100 supply shock in any given six-month window has been material through 2024 and 2025. The cost of an unmanaged supply shock is measured in engineering days and customer-facing SLA degradation. The cost of building portability is measured in engineering days paid once, in advance.
The teams that absorbed this cost upfront are the ones operating quietly when the rest of the ML infrastructure community is filing support tickets about spot availability queues. It is not a flashy engineering problem, but it is one of the few infrastructure investments that pays off with certainty rather than probability.
Design your inference stack to survive any accelerator supply cycle
ZML routes across available hardware automatically. When one vendor is constrained, traffic shifts to another.