The GPU price is visible on the invoice. The hidden cost is everything that happens when that GPU stops being available at the price, availability, or performance level you planned around. We have been tracking this cost since ZML's early-access program launched, and the numbers are consistently larger than teams estimate before they hit the constraint.
This post quantifies the hidden costs we observed, describes the failure modes where they surface, and explains why the cost is systematically underestimated until it is paid.
The Three Categories of Hidden Cost
Accelerator vendor lock-in creates costs in three distinct categories, each with different timing and visibility.
The first category is migration cost: the engineering time required to move a serving stack to a different hardware type. This is the most visible category but is consistently underestimated. Teams estimate migration cost based on the surface area of their inference call, not the surface area of their entire serving stack. The inference call is one function call. The serving stack has containers, custom kernels, quantization pipelines, monitoring integrations, and load testing infrastructure, all of which have hardware dependencies that only surface during migration.
The second category is supply constraint cost: the cost paid during a hardware unavailability event when migration has not been completed. This includes elevated spot prices when spot availability drops (on-demand premiums of 2 to 3x are common during constraint periods), degraded throughput while operating below intended capacity, and the engineering time spent managing the supply event rather than building product.
The third category is opportunity cost: capabilities foregone because the stack was not designed for hardware flexibility. Teams that cannot easily route to MI300X cannot take advantage of its cost advantage for batch workloads. Teams that cannot route to on-demand H100 when their latency SLA is at risk have fewer levers to pull during traffic spikes. These costs are invisible in any accounting that looks only at actual spend, but they are real.
Migration Cost: What We Actually Measured
Across teams in our early-access program that undertook H100-to-MI300X migrations, we tracked engineering time from migration start to production validation. Migration start was defined as the first engineering commit specifically targeting AMD hardware. Production validation was defined as stable traffic on MI300X with no outstanding latency or reliability regressions from pre-migration baseline.
The median migration time for a serving stack with what we categorize as moderate customization (one to two custom CUDA kernels, TensorRT-based quantization, standard monitoring with some nvidia-smi integration) was 2.6 weeks of engineering time spread over 4 to 5 calendar weeks. Teams with high customization (custom attention kernels, custom batching logic, GPU-specific memory management) saw migration times of 5 to 8 weeks.
Teams with minimal customization (standard vLLM or TGI, no custom kernels, generic container tooling) completed migrations in 4 to 7 days. This is the group that had designed their stack closest to framework abstraction boundaries, and their migration cost was proportionally lower.
These numbers are for intentional, planned migrations. Emergency migrations under supply pressure, where the team had not done the preparation work, took 30 to 50 percent longer due to the combination of time pressure, incomplete documentation of hardware dependencies, and the need to validate under live traffic constraints simultaneously.
Supply Constraint Cost: A Composite Estimate
Quantifying supply constraint cost requires combining several components. Hardware premium: H100 on-demand pricing is typically 2.2 to 2.8x spot pricing in comparable cloud regions. A team that would have been running on $2.50/hr H100 spot pays $5.50 to $7.00/hr on-demand during a spot constraint event. For a serving deployment running 16 H100 nodes, a 10-day constraint event represents $30,000 to $70,000 in incremental hardware cost above planned spend.
Capacity degradation: when spot availability drops and on-demand economics are not acceptable, teams often reduce serving capacity rather than pay the premium. Reduced capacity means queue depth increases, p99 latency degrades, and the team begins receiving SLA alerts. The engineering time responding to these alerts is real work cost that has no direct line item in any budget but displaces feature development and other infrastructure work.
For teams that have built MI300X as a fallback, the supply constraint event costs a policy update (minutes) and the small performance delta of running on secondary hardware. For teams without a fallback, it costs the on-demand premium or the capacity degradation. The difference is a one-time upfront investment in multi-hardware support versus recurring exposure to supply constraint events.
Why Migration Cost Is Systematically Underestimated
The underestimation is not carelessness. It follows from a rational information problem: the full hardware dependency surface of a serving stack is not visible until you try to move it. Application engineers who own the inference call see a function call. They do not own the container toolchain, the monitoring pipeline, or the custom kernel that a previous team member contributed eighteen months ago. Each of those is a migration cost that surfaces only when the migration begins.
There is also a retrospective discount effect: after a migration is complete, the work that was done looks smaller than it felt during execution. Teams that completed a migration in 3 weeks tend to estimate future migrations at 1 to 2 weeks when asked in retrospect, because the disruptive parts (the debugging sessions, the false starts, the unexpected kernel incompatibility) are not salient after resolution. This means estimates made from retrospective reference points are systematically optimistic.
Reducing the Exposure Without a Full Rewrite
The goal is not to rewrite all existing serving infrastructure to be hardware-agnostic. The goal is to reduce the migration surface for the parts of the stack that change most often, and to build the dispatch layer in a way that does not require touching application code when hardware changes.
The dispatch layer is where the leverage lives. If routing a request to H100 versus MI300X is a policy evaluation rather than a code path, the migration cost for adding a new hardware family is a configuration change, not an engineering project. The container toolchain, the quantization pipeline, and the monitoring integration still require hardware-specific work, but that work can be amortized across all traffic that routes to the new hardware, rather than paid per service that needs to support it.
This is the architectural argument behind ZML. We are not claiming to eliminate migration cost. We are saying that the dispatch layer is the highest-leverage place to invest in hardware flexibility, and that investment pays dividends every time a supply constraint event occurs, every time a new accelerator generation ships, and every time a cost-per-token optimization requires changing hardware mix.
Build vendor independence into your inference stack from day one
ZML hardware-agnostic API means switching accelerators is a config change, not an engineering project.