Engineering articles from ZML
Benchmarks, infrastructure guides, and technical deep-dives on ML inference, accelerator selection, and production serving from the ZML team.
Running LLM Inference on AMD MI300X: What Infrastructure Teams Need to Know
AMD MI300X offers compelling cost-per-token for large batch workloads, but the path from NVIDIA to AMD is rarely smooth. We document what actually changed when we moved Llama 3 70B workloads across to MI300X hardware.
Read article
How to Choose an Accelerator for Production LLM Serving in 2026
H100, MI300X, Gaudi2, or something else? The decision used to be simple. This guide covers the criteria that actually matter when spot supply fluctuates and your cost model has to hold.
GPU Supply Volatility Is Not Going Away. Here Is How to Design Around It
H100 spot availability has dipped below 60 percent in multiple quarters. Teams that built for a single accelerator vendor are paying the price in engineering hours and idle capacity.
Cost Optimization for Production LLM Inference: A Practical Framework
Inference cost is now the biggest line item for teams running LLMs in production. We break down the three levers that actually move the needle: hardware mix, batching strategy, and routing policy.
How ZML Routes an Inference Request: A Walkthrough
A request comes in. ZML has 200ms to decide which accelerator node handles it. This post walks through every step of the routing decision, from policy evaluation to dispatch.
H100 vs MI300X on Llama 3: A Head-to-Head Benchmark
We ran identical Llama 3 8B and 70B workloads on H100 SXM and MI300X nodes across 5 batch sizes and 3 precision formats. Here are the raw numbers from our internal test cluster.
ZML Python SDK 0.3: Policy Chains, Async Support, and Better Error Messages
Version 0.3 of the ZML Python SDK ships policy chain configuration, native async/await support, and significantly improved error messages when accelerator routing fails.
Cold Start Latency Across Heterogeneous Accelerators: Measurements and Mitigations
Cold start on AMD MI300X can be 40 to 60 percent longer than H100 for the same model at comparable precision. We explain why, and how ZML warmup scheduling reduces the impact.
Serving Vision Models Across Multiple Accelerators Without Rewriting Your Pipeline
Vision models like CLIP and Stable Diffusion have different accelerator affinity than LLMs. We explain how ZML handles mixed-model fleets where text and vision workloads compete for the same hardware pool.
Model Serving Infrastructure Across Hardware Generations: What Changes and What Does Not
From V100 to A100 to H100 to MI300X, the serving stack fundamentals have stayed surprisingly stable while the dispatch layer has gotten increasingly complex.
The Hidden Cost of Accelerator Vendor Lock-In
The visible cost is the GPU price. The invisible cost is the 3-week migration sprint every time supply changes or a new accelerator generation ships. We tracked this cost across teams in our early-access program.
Why We Built ZML: Inference Should Be Hardware-Agnostic by Default
The first time we had to rewrite a production inference pipeline because H100s were backordered, we started thinking there had to be a better design. This is the story behind ZML.