Sign In Get Started
Docs

API Reference

Complete reference for the ZML Python SDK. The SDK communicates with your ZML dispatcher over gRPC or REST. All parameters and return types are documented below.

zml.Client

The entry point for all ZML operations. Instantiating a client establishes a connection pool to your ZML dispatcher endpoint.

zml.Client(
    endpoint=None,        # str, defaults to ZML_ENDPOINT env var
    api_key=None,         # str, defaults to ZML_API_KEY env var
    timeout=30,           # int, seconds per request
    max_connections=10,  # int, connection pool size
    config_path=None       # str, path to zml.toml (auto-discovered if None)
)

If endpoint and api_key are both omitted, the client reads from environment variables and then from zml.toml in the current working directory (or any parent directory).

client.infer()

Submit an inference request to the ZML dispatcher. The dispatcher selects an accelerator node according to the active routing policy and returns the model output with metadata.

client.infer(
    model,               # str, required: model identifier
    prompt,              # str, required: input text
    max_tokens=512,      # int, max output tokens
    temperature=0.7,     # float, sampling temperature
    routing=None,        # str, override default policy for this call
    accelerator_id=None, # str, pin to a specific node (bypasses routing)
    stream=False          # bool, if True returns an async generator
) -> InferResult

The InferResult object has two fields:

  • result.text: the generated text output
  • result.metadata: an InferMeta object (see below)

InferMeta fields

Field Type Description
accelerator_id str ID of the accelerator node that served this request
routing_policy str The routing policy applied to this request
time_to_first_token float Seconds from request submission to first output token
queue_depth_at_dispatch int Queue depth on the chosen node at dispatch time
total_tokens int Prompt tokens + generated tokens

Routing Policy Schema

Routing policies are defined in zml.toml under the [policy] section, or passed per-call via the routing parameter.

Built-in policies

Policy Description
latency Route to the node with the lowest current time-to-first-token estimate
cost Route to the node with the lowest cost_weight that is below queue threshold
availability Route to the node with the most remaining capacity. Maximizes utilization.
round_robin Distribute requests evenly across all healthy nodes regardless of load

Custom policy example

[policy.custom_cost_latency]
strategy = "weighted"
latency_weight = 0.4
cost_weight = 0.6
queue_threshold = 8    # fallback to latency if queue > 8

Accelerator IDs

Use these type identifiers in your zml.toml accelerator blocks. ZML probes each node at startup and validates the reported hardware against the declared type.

Type ID Hardware Min driver Notes
nvidia_h100 NVIDIA H100 SXM / PCIe 525.85 CUDA 12.0+
nvidia_a100 NVIDIA A100 SXM / PCIe 515.48 CUDA 11.7+
amd_mi300x AMD Instinct MI300X ROCm 6.0 See cold-start note below
intel_gaudi2 Intel Gaudi2 1.12 Habana Labs driver
aws_inferentia2 AWS Inferentia2 Neuron SDK 2.18 inf2.xlarge and above
AMD MI300X cold start: Initial model loading on MI300X can run 40 to 60 percent longer than on H100 for the same model at comparable precision. ZML's warmup scheduler pre-loads models during low-traffic windows to mitigate this. See the blog post on cold start latency mitigations for details.

Error Codes

ZML raises zml.ZMLError with a code attribute on all non-success responses:

Code Meaning
NO_HEALTHY_NODES All registered accelerator nodes failed health checks. Check connectivity and driver status.
MODEL_NOT_LOADED The requested model is not loaded on any available node. Run zml serve for this model first.
QUEUE_FULL All nodes for this model are over the queue threshold. Retry with backoff or scale out nodes.
AUTH_FAILED Invalid or missing API key. Check ZML_API_KEY environment variable.
TIMEOUT Request exceeded the configured timeout value. Increase timeout or reduce prompt length.

Rate Limits

Rate limits are enforced per API key at the ZML dispatcher level. Default limits for the managed service:

Plan Requests/min Concurrent
Maker 60 5
Team 600 50
Production Custom Custom

Self-hosted deployments have no rate limits imposed by ZML. Limits can be configured at the reverse proxy layer independently.