The first time we rewrote a production inference pipeline because H100s went on backorder, we spent three weeks on it. The model had not changed. The traffic had not changed. The performance requirements had not changed. We were doing the same work at the end of three weeks that we had been doing at the start, except now on different hardware. That was the first time the question occurred clearly: why is hardware type a design constraint rather than a runtime decision?
It took two more of those events before we started building ZML. This is that story.
The Problem Is Not the GPU Price
The GPU price is visible and people argue about it constantly. The problem that motivated ZML is not the GPU price; it is the implicit assumption in most serving stacks that hardware type is fixed at deployment time.
When you build your inference serving stack with H100 as a fixed assumption, that assumption propagates through the stack in ways that are not immediately obvious. The container base image is NVIDIA. The custom attention kernels are CUDA compiled. The monitoring integration calls nvidia-smi. The quantization pipeline generates TensorRT engines for H100. Each of these is a reasonable engineering choice given a single-hardware deployment. Each one becomes a migration cost when the hardware changes.
The serving frameworks abstract some of this. vLLM, TGI, and similar tools sit above the hardware layer and provide a consistent inference API across hardware families where they have backend support. But the dispatch layer, the part of the system that decides which physical node handles a given request, is almost never included in that abstraction. The dispatch layer in most production serving systems is either implicit (load balancer round-robin across identical nodes) or hardcoded (routing logic that assumes specific hardware types).
When hardware changes, the dispatch layer is what breaks. It is also what most teams are least prepared to change quickly, because it touches both the serving framework configuration and the infrastructure layer, and it typically has undocumented assumptions that only surface under load.
What We Observed in Teams Around Us
Before building ZML, we talked with engineers from teams running LLMs in production in the early-stage AI application ecosystem in 2024. A pattern appeared in almost every conversation. Every team had experienced at least one supply constraint event where their primary hardware type became unavailable or unaffordably expensive at spot pricing. Every team had spent engineering time on that event doing work that felt like it should not exist: debugging why the same model ran differently on different hardware, updating infrastructure configuration, managing traffic manually during the transition period.
The teams that had handled it most cleanly were not the ones with the most sophisticated infrastructure. They were the ones whose serving framework sat clearly above the hardware layer, with explicit, documented assumptions about where hardware-specific code lived. Those teams were still paying migration costs, but they were paying them in hours, not weeks.
The teams that had handled it worst were the ones whose serving code had grown hardware assumptions through organic accretion: the CUDA-specific kernel added six months ago for a performance improvement, the monitoring script that calls nvidia-smi because the engineer who wrote it only knew nvidia-smi, the batch size configuration committed with H100 memory limits hardcoded in a comment. Each of these was added rationally in context. Together they made the dispatch layer opaque and the migration slow.
The Design Principle Behind ZML
The principle we started with: hardware type should be a runtime routing decision, not a deployment-time design constraint. What we mean by this precisely: the serving code that your application writes should not contain knowledge of which physical accelerator handles the request. That knowledge should live in a routing policy, expressed as configuration, evaluated at request dispatch time.
This is not a novel idea in distributed systems. Load balancers have routed to heterogeneous server pools by capability for decades. DNS-based routing sends traffic to different endpoints based on geography or health. Session-aware routing has matched clients to backends based on affinity rules for as long as web applications have existed. The concept of abstracting the target selection from the application is well-established.
What was missing for ML inference was a routing system that understood the specific heterogeneity of accelerator hardware: different memory pools, different throughput profiles at different batch sizes, different precision format support, different cold start characteristics. A generic load balancer does not make routing decisions based on "H100 wins for this request because its p95 latency target is 200ms and MI300X is currently at 80 percent queue capacity." That decision requires inference-specific understanding of both hardware and workload, and it needs to happen in milliseconds.
ZML Is Not a Framework. It Is a Dispatch Layer
We want to be precise about what ZML is and is not, because the boundaries matter for understanding what problem it solves.
ZML is not a serving framework. We do not compete with vLLM, TGI, or TensorRT-LLM. Those frameworks run the actual inference on the accelerator nodes. ZML sits above them and decides which node a request goes to. The inference call passes through ZML and reaches whatever framework is running on the selected node.
ZML is not a monitoring or observability tool, though it emits routing metadata that is useful for observability. It is not a model registry or model versioning system. It is not a serving framework configuration manager.
ZML is a dispatch layer with a routing policy language. You define policies that express your hardware preferences, latency targets, and cost priorities. ZML evaluates those policies against live accelerator state and dispatches requests accordingly. When hardware availability changes, you update policy configuration. Your application code does not change.
Where We Are Now
ZML is running on real production serving infrastructure for teams in our early-access program. The current scope is H100, MI300X, and Gaudi2 routing via the Python SDK and a REST API. Policy chains, async support, and the routing metadata API are all shipped and used in production.
We are an independently funded team of three in Paris. We built this because we believed the dispatch problem was real and was not going to be solved by the serving frameworks that had a different scope. The teams we have worked with in early access have confirmed that the dispatch layer is where the migration cost lives, and that separating it from the serving layer is the right architectural move for teams that expect to operate across hardware generations.
If you are running LLM inference in production and have felt the weight of a hardware change in your serving stack, that experience is the exact problem we designed ZML to prevent. The quickstart is at zmlai.org/docs/quickstart.html and the free Maker plan covers single-node evaluation without a credit card. We would rather have you form your own opinion than take our word for it.
See what we built and try it yourself
ZML is live and free to start. Connect your first accelerator pool in minutes with the Python SDK.