Sign In Get Started
Back to Blog

How ZML Routes an Inference Request: A Walkthrough

A request comes in. ZML has 200ms to decide which accelerator node handles it. This post walks through every step of the routing decision, from policy evaluation to dispatch.

Diagram illustrating ZML inference request routing flow from client through dispatcher to accelerator nodes

When a request arrives at ZML, there is a window of roughly 200 milliseconds before the system needs to have committed that request to a specific accelerator node. In practice, the decision happens in about 8 to 12 milliseconds for the typical case. This post walks through every step of that decision path, from the moment the request reaches the ZML gateway to the moment it is handed off to a compute node for execution.

The goal is not to describe an abstract architecture. It is to show the specific reasoning at each decision point so that ML infrastructure engineers can understand what ZML is optimizing for and what happens when the inputs to that optimization change.

Step 1: Request Intake and Policy Resolution

The inference request arrives at the ZML gateway as a standard HTTP/2 call with a model identifier, a messages payload, and an optional accelerator hint parameter. The first thing the gateway does is resolve the routing policy for this request.

A routing policy is a named configuration object that defines an ordered preference chain for hardware types, a latency target in milliseconds, and a cost priority level (values: cost, balanced, latency). Policies are attached to model registrations, but they can be overridden at call time via the policy parameter. The policy lookup is a key-value read from an in-memory policy store, targeting sub-millisecond latency. Policy configuration is covered in the API reference, but a typical policy object looks like this:

name: "llama3-70b-batch"
accelerator_preference:
  - family: MI300X
    priority: 1
  - family: H100
    priority: 2
latency_target_ms: 2000
cost_priority: cost

If no explicit policy is passed and no model-level default is configured, ZML falls back to the account-level default policy. The account default is set to cost_priority: balanced with a preference order that reflects your registered accelerator families in registration order. Most teams configure explicit policies per model class and rely on the default only for development traffic.

Step 2: Accelerator Availability Query

With the policy resolved, the router queries the live accelerator state for each hardware family in the preference chain. The accelerator state table is a lightweight in-memory structure updated by a background health aggregator that polls each registered node every 5 seconds. The state per node includes: current queue depth, free memory (GiB), last successful inference timestamp, and a derived health score (0 to 100, computed from queue depth and recent error rate).

The availability query does not block on the network. It reads from the local state table. This is why routing latency stays in the single-digit milliseconds range even when you have a large node fleet. The tradeoff is that the state table has a maximum staleness of 5 seconds. In practice, a node that fails between state updates will be caught by the post-dispatch health check rather than the routing query, and the request is re-dispatched automatically.

For each accelerator family in the policy preference chain, the availability query returns a ranked list of nodes sorted by effective throughput capacity, defined as the ratio of current queue depth to historical tokens-per-second throughput. A node with low queue depth and high historical throughput ranks first; a node near capacity ranks last.

Step 3: Cost Evaluation

When the policy's cost_priority is cost or balanced, the router evaluates a cost estimate for each candidate hardware family. Cost data comes from a provider cost table that is updated periodically with spot prices for each registered cloud account. For on-premise hardware, cost entries are user-defined at a normalized cost-per-GPU-hour rate.

The cost evaluation is not a simple minimum-price selection. It is weighted by the policy's latency target. If the cheapest option has a queue depth that would push estimated response time beyond the latency target, the router considers the next option in the preference chain even if it is more expensive. This prevents a common failure mode where aggressive cost routing sends traffic to an overloaded cheap node and misses SLA.

For latency cost priority, the cost evaluation step is skipped entirely and the router selects the first available node in the preference chain that has acceptable queue depth regardless of price. This mode is appropriate for synchronous user-facing workloads where latency SLA is fixed and cost is secondary.

Step 4: Dispatch and Acknowledgment

With a target node selected, the router dispatches the request to the node's inference endpoint. The dispatch is a direct HTTP/2 call to the serving framework running on the selected node (vLLM, TGI, or a custom server). ZML does not rewrite the request payload; it adds a routing metadata header that the serving framework can optionally log for observability.

The routing metadata header contains: the selected node ID, the policy name that was applied, the accelerator family, the estimated queue depth at dispatch time, and a routing decision ID for trace correlation. This header is preserved in the response and surfaced in the ZML Python SDK's response object as response.routing.

The acknowledgment from the inference node triggers a state table update: the dispatched request increments the node's queue depth counter immediately, before the inference actually runs. This prevents the race condition where multiple routing decisions happen simultaneously and all see the same queue depth snapshot for a node, causing over-assignment.

Step 5: Failure Handling and Fallback

If the dispatch call fails with a network error or the inference node returns a 5xx response, ZML re-initiates the routing decision from step 2. The failed node is temporarily removed from the candidate list for a configurable cooldown period (default: 30 seconds). The re-dispatch selects the next-best available node. If all nodes in the preferred accelerator family are unavailable, the router moves to the next family in the preference chain.

One scenario worth understanding explicitly: if all nodes across all configured accelerator families are unavailable or over capacity, ZML returns a 503 to the caller with a body indicating queue saturation. It does not hold the request in queue indefinitely by default. Queue-based waiting is a configuration option per policy (queue_on_saturation: true with a max_queue_ms limit). The default is to fail fast and let the caller retry, which is appropriate for synchronous API use cases. Batch job policies typically enable queue-based waiting.

What Changes When You Add a New Accelerator Family

Adding an AMD MI300X fleet to a cluster that previously ran only H100 requires three configuration changes: register the MI300X nodes with their endpoint addresses and cost rates, update relevant routing policies to include MI300X in the preference chain at the desired priority level, and validate that the models you serve are configured in the serving framework with ROCm support on those nodes.

ZML does not require any changes to the inference call itself. The same client.infer() call that was routing to H100 will start routing to MI300X for requests that match the updated policy. The serving API is hardware-agnostic by design. This is the specific guarantee that makes multi-accelerator fleets operationally manageable: you tune policy configuration, not inference code.

Configure routing policies for your accelerator pool

ZML routing policies give you latency targets, cost ceilings, and hardware affinity in one config object.

Policy reference Get started