The hardware generations that ML teams have navigated from 2018 to 2025 span V100, A100, H100, and now MI300X. Each transition was sold as an upgrade with a relatively clean migration path. In practice, each one required teams to rebuild specific parts of their serving stack while other parts survived unchanged. The pattern of what changes and what does not is informative, because understanding it helps predict what the next generation will require.
This post maps each hardware generation against the serving stack, identifies what each transition actually broke, and looks at what has accumulated as permanent complexity in the dispatch layer over the years.
The V100 Era: When the Stack Was Simple
V100 marked the transition from research-grade GPU serving to early production deployment. The characteristic of V100-era serving stacks is simplicity: most production serving was single-GPU, batch sizes were small by current standards, and the complexity surface was limited to model loading, basic batching, and CUDA runtime management.
The serving stack in 2018-2020 was typically a Python process with a model loaded into GPU memory, a simple request queue, and a synchronous forward pass per request. This model worked at the traffic volumes and model sizes of the time. The V100 had 32GB HBM2, which was sufficient for the model sizes in production use.
The dispatch layer, such as it was, had no hardware selection logic because there was no choice to make. You had the nodes you had, and you ran inference on them. The concept of routing a request to a specific hardware type did not exist in most serving stacks.
A100 Transition: The First Forced Rebuild
The A100 introduced 80GB HBM2e (SXM4 variant), which was transformative for model size but created the first significant migration surface. Teams that had been running serving on V100 with model weights that used the full 32GB available had to rethink their memory management when moving to a machine with 80GB, because the assumptions built into their batching code (how many sequences could coexist in memory, what KV-cache size was reasonable) had changed.
More consequentially, A100 introduced tensor parallelism as a first-class serving pattern for large models. Models that required multiple V100s via pipeline parallelism could now use NVLink-based tensor parallelism across A100 SXM nodes with much lower communication latency. Frameworks that supported TP were mostly research tools in 2020. By 2022, production teams running models above 13B parameters needed TP, and the frameworks (FasterTransformer, early vLLM, DeepSpeed-Inference) had varying levels of production readiness.
The A100 transition also introduced the first explicit hardware-specific compilation paths. TensorRT became a production requirement for teams squeezing maximum throughput on A100. CUDA kernels tuned for A100's tensor core layout performed significantly better than V100-tuned equivalents, and teams running untouched V100-era kernels on A100 left substantial throughput on the table.
H100 Transition: FP8 and the Attention Kernel Arms Race
The H100 generation is the most significant kernel architecture change since the initial CUDA GPU compute era. Two changes drove the majority of the migration work: FP8 support via the Hopper FP8 tensor cores, and the introduction of FlashAttention-2 and its H100-specific optimizations.
FP8 support was not a drop-in benefit. Teams using TensorRT or custom quantization needed to rebuild quantization pipelines for FP8 E4M3/E5M2, validate output quality, and handle the cases where FP8 precision was insufficient for specific model architectures. The frameworks (vLLM FP8 support, TRT-LLM FP8) came online incrementally through 2023-2024. Teams that moved to H100 and expected FP8 gains immediately often ran BF16 for months while FP8 kernel support matured in their stack.
FlashAttention-2 on H100 changed the performance envelope for attention-heavy models. Teams that had custom attention implementations not using FlashAttention had to migrate to benefit from H100's attention optimization. This was a meaningful code change for serving stacks with heavily customized attention kernels.
MI300X: The First Cross-Vendor Transition
Each previous generation transition was CUDA to CUDA. The MI300X generation is the first time production serving teams have faced a real choice between vendor toolchains. The migration surface is fundamentally different: it is not "tune existing code for new hardware" but "port code from one computing stack to another."
The container toolchain changes. The kernel compilation paths change. Monitoring instrumentation changes (nvidia-smi to rocm-smi, different metric namespaces). Existing TensorRT pipelines require replacement with alternatives in the ROCm ecosystem. Custom CUDA kernels need HIP translation or replacement with ROCm equivalents.
What survives the cross-vendor transition is the application layer and the serving framework interface. If your serving code calls vLLM through its Python API, that code is portable. If your serving code has CUDA-specific calls anywhere below the serving framework, that code is not portable. The architectural lesson from observing multiple cross-vendor migrations is that the serving framework abstraction level is where portability lives. Code that bypasses the framework for performance reasons pays a portability cost that becomes material during the vendor transition.
What Has Not Changed Across Generations
The core serving primitives have been remarkably stable. Batching theory has not changed since the A100 era. The tradeoff between throughput and latency through batch size is the same on MI300X as it was on V100. KV-cache management principles (PagedAttention is an implementation, not a fundamental new concept) are derived from operating system page table design, not GPU-specific innovation. Request queuing, health checking, and model versioning strategies are unchanged.
What has changed is the dispatch complexity. V100 serving had no dispatch problem: one hardware type, one backend, one runtime. A100 serving introduced TP configuration complexity. H100 added FP8 precision routing decisions. MI300X adds cross-vendor backend selection. The dispatch layer has accumulated complexity with each generation, and none of that complexity was removed when the new hardware arrived. It was added on top.
The implication for teams building serving infrastructure today: the dispatch layer deserves architectural investment proportional to its complexity. Hardcoding hardware assumptions at the dispatch layer means rebuilding it with each hardware generation. Building a routing abstraction at the dispatch layer means each new hardware generation is an additive configuration change, not a rebuild. That is the specific problem that motivated ZML, and the cross-vendor MI300X transition is the generation where the cost of not investing in dispatch abstraction became measurable in engineering weeks, not hours.
Build a serving layer that absorbs the next hardware generation without a rewrite
ZML abstracts the hardware so your inference code does not change when your accelerator fleet does.