Run your models on any accelerator, not just the one you got locked into
ZML serves production ML inference across GPUs and other accelerators from a single stack, so your infrastructure team can choose hardware on price and supply instead of whatever your framework happens to support.
Python SDK
import zml
# Connect to your ZML deployment
client = zml.Client(endpoint="https://api.yourinfra.com")
# Route inference to best available accelerator
response = client.generate(
model="llama3-70b",
prompt="Summarize the quarterly report",
routing="latency" # or "cost", "availability"
)
print(response.text, response.accelerator_used)
# "The Q3 results show..." | "nvidia-h100-node-3"
Runs on your hardware
One stack. Every accelerator.
Deploy ZML and let your team pick hardware based on what is available and affordable, not what your framework supports.
Latency-aware routing
ZML measures queue depth and time-to-first-token across every node in real time, routing each request to the fastest available accelerator at that moment.
Unified model API
One gRPC or REST endpoint serves all your models regardless of which accelerator is running them. No changes to your application code when hardware changes.
Accelerator abstraction
Write model configs once. ZML compiles and serves them on H100, A100, MI300X, Gaudi2, and other supported targets without per-device code paths.
Live observability
Prometheus metrics and OpenTelemetry traces ship with every deployment. See per-node throughput, latency percentiles, and routing decisions in your existing dashboards.
Cost routing
Configure per-accelerator cost weights and ZML will route bulk or background workloads toward cheaper nodes while keeping interactive requests on low-latency hardware.
Data stays on-prem
ZML runs entirely inside your infrastructure. No inference traffic leaves your network. Bring your own VPC, air-gapped data center, or co-location facility.
Your accelerators. Your data center.
ZML deploys as a set of lightweight services that sit in front of your existing accelerator fleet, adding routing and observability without replacing your provisioning workflow.
From model to production in three steps
ZML is built to fit into existing infrastructure workflows, not replace them.
Register your accelerators
Point ZML at each node with a YAML config or Terraform module. It probes available memory, compute capability, and current load automatically. No agent required on the host.
Deploy your model
Run zml serve model.yaml and ZML compiles the model for each target architecture in your pool and loads it onto available hardware. Hot-swap supported with zero downtime.
Call a single endpoint
Your application sends inference requests to ZML's unified API. Routing decisions happen in the dispatcher under 2 ms overhead. Add hardware later without touching application code.
Stop picking hardware based on what your stack allows
ZML takes the accelerator out of the critical path so your infrastructure team can optimize for price, supply, and performance independently.