Sign In Get Started
Blog

Engineering articles from ZML

Benchmarks, infrastructure guides, and technical deep-dives on ML inference, accelerator selection, and production serving from the ZML team.

AMD MI300X accelerator in a server rack
Server comparison of GPU accelerator options
Infrastructure

How to Choose an Accelerator for Production LLM Serving in 2026

H100, MI300X, Gaudi2, or something else? The decision used to be simple. This guide covers the criteria that actually matter when spot supply fluctuates and your cost model has to hold.

Ana Petrov
GPU supply chain analysis chart
Strategy

GPU Supply Volatility Is Not Going Away. Here Is How to Design Around It

H100 spot availability has dipped below 60 percent in multiple quarters. Teams that built for a single accelerator vendor are paying the price in engineering hours and idle capacity.

Steeve Morin
Cost optimization dashboard for inference workloads
Engineering

Cost Optimization for Production LLM Inference: A Practical Framework

Inference cost is now the biggest line item for teams running LLMs in production. We break down the three levers that actually move the needle: hardware mix, batching strategy, and routing policy.

Ana Petrov
Diagram of ZML inference routing logic
Engineering

How ZML Routes an Inference Request: A Walkthrough

A request comes in. ZML has 200ms to decide which accelerator node handles it. This post walks through every step of the routing decision, from policy evaluation to dispatch.

Marcus Webb
Benchmark comparison chart H100 vs MI300X
Research

H100 vs MI300X on Llama 3: A Head-to-Head Benchmark

We ran identical Llama 3 8B and 70B workloads on H100 SXM and MI300X nodes across 5 batch sizes and 3 precision formats. Here are the raw numbers from our internal test cluster.

Marcus Webb
Python SDK release announcement
Release

ZML Python SDK 0.3: Policy Chains, Async Support, and Better Error Messages

Version 0.3 of the ZML Python SDK ships policy chain configuration, native async/await support, and significantly improved error messages when accelerator routing fails.

Ana Petrov
Latency measurements across accelerator types
Engineering

Cold Start Latency Across Heterogeneous Accelerators: Measurements and Mitigations

Cold start on AMD MI300X can be 40 to 60 percent longer than H100 for the same model at comparable precision. We explain why, and how ZML warmup scheduling reduces the impact.

Ana Petrov
Vision model serving across mixed GPU hardware
Engineering

Serving Vision Models Across Multiple Accelerators Without Rewriting Your Pipeline

Vision models like CLIP and Stable Diffusion have different accelerator affinity than LLMs. We explain how ZML handles mixed-model fleets where text and vision workloads compete for the same hardware pool.

Marcus Webb
Timeline of hardware generations for ML serving
Infrastructure

Model Serving Infrastructure Across Hardware Generations: What Changes and What Does Not

From V100 to A100 to H100 to MI300X, the serving stack fundamentals have stayed surprisingly stable while the dispatch layer has gotten increasingly complex.

Steeve Morin
Cost analysis of accelerator vendor lock-in
Strategy

The Hidden Cost of Accelerator Vendor Lock-In

The visible cost is the GPU price. The invisible cost is the 3-week migration sprint every time supply changes or a new accelerator generation ships. We tracked this cost across teams in our early-access program.

Steeve Morin
ZML founders discussing inference orchestration design
Company

Why We Built ZML: Inference Should Be Hardware-Agnostic by Default

The first time we had to rewrite a production inference pipeline because H100s were backordered, we started thinking there had to be a better design. This is the story behind ZML.

Steeve Morin