Sign In Get Started
ML Inference Orchestration

Run your models on any accelerator, not just the one you got locked into

ZML serves production ML inference across GPUs and other accelerators from a single stack, so your infrastructure team can choose hardware on price and supply instead of whatever your framework happens to support.

NVIDIA H100 / A100 AMD MI300X Intel Gaudi2 ZML Dispatcher latency-aware routing Output gRPC / REST

Python SDK

Python
import zml

# Connect to your ZML deployment
client = zml.Client(endpoint="https://api.yourinfra.com")

# Route inference to best available accelerator
response = client.generate(
    model="llama3-70b",
    prompt="Summarize the quarterly report",
    routing="latency"  # or "cost", "availability"
)

print(response.text, response.accelerator_used)
# "The Q3 results show..." | "nvidia-h100-node-3"

Runs on your hardware

Platform

One stack. Every accelerator.

Deploy ZML and let your team pick hardware based on what is available and affordable, not what your framework supports.

Latency-aware routing

ZML measures queue depth and time-to-first-token across every node in real time, routing each request to the fastest available accelerator at that moment.

Unified model API

One gRPC or REST endpoint serves all your models regardless of which accelerator is running them. No changes to your application code when hardware changes.

Accelerator abstraction

Write model configs once. ZML compiles and serves them on H100, A100, MI300X, Gaudi2, and other supported targets without per-device code paths.

Live observability

Prometheus metrics and OpenTelemetry traces ship with every deployment. See per-node throughput, latency percentiles, and routing decisions in your existing dashboards.

Cost routing

Configure per-accelerator cost weights and ZML will route bulk or background workloads toward cheaper nodes while keeping interactive requests on low-latency hardware.

Data stays on-prem

ZML runs entirely inside your infrastructure. No inference traffic leaves your network. Bring your own VPC, air-gapped data center, or co-location facility.

Infrastructure

Your accelerators. Your data center.

ZML deploys as a set of lightweight services that sit in front of your existing accelerator fleet, adding routing and observability without replacing your provisioning workflow.

Server racks in a production datacenter running mixed GPU accelerators
How it works

From model to production in three steps

ZML is built to fit into existing infrastructure workflows, not replace them.

01

Register your accelerators

Point ZML at each node with a YAML config or Terraform module. It probes available memory, compute capability, and current load automatically. No agent required on the host.

02

Deploy your model

Run zml serve model.yaml and ZML compiles the model for each target architecture in your pool and loads it onto available hardware. Hot-swap supported with zero downtime.

03

Call a single endpoint

Your application sends inference requests to ZML's unified API. Routing decisions happen in the dispatcher under 2 ms overhead. Add hardware later without touching application code.

Stop picking hardware based on what your stack allows

ZML takes the accelerator out of the critical path so your infrastructure team can optimize for price, supply, and performance independently.