API Reference
Complete reference for the ZML Python SDK. The SDK communicates with your ZML dispatcher over gRPC or REST. All parameters and return types are documented below.
zml.Client
The entry point for all ZML operations. Instantiating a client establishes a connection pool to your ZML dispatcher endpoint.
zml.Client(
endpoint=None, # str, defaults to ZML_ENDPOINT env var
api_key=None, # str, defaults to ZML_API_KEY env var
timeout=30, # int, seconds per request
max_connections=10, # int, connection pool size
config_path=None # str, path to zml.toml (auto-discovered if None)
)
If endpoint and api_key are both omitted, the client reads from environment variables and then from zml.toml in the current working directory (or any parent directory).
client.infer()
Submit an inference request to the ZML dispatcher. The dispatcher selects an accelerator node according to the active routing policy and returns the model output with metadata.
client.infer(
model, # str, required: model identifier
prompt, # str, required: input text
max_tokens=512, # int, max output tokens
temperature=0.7, # float, sampling temperature
routing=None, # str, override default policy for this call
accelerator_id=None, # str, pin to a specific node (bypasses routing)
stream=False # bool, if True returns an async generator
) -> InferResult
The InferResult object has two fields:
result.text: the generated text outputresult.metadata: anInferMetaobject (see below)
InferMeta fields
| Field | Type | Description |
|---|---|---|
accelerator_id |
str | ID of the accelerator node that served this request |
routing_policy |
str | The routing policy applied to this request |
time_to_first_token |
float | Seconds from request submission to first output token |
queue_depth_at_dispatch |
int | Queue depth on the chosen node at dispatch time |
total_tokens |
int | Prompt tokens + generated tokens |
Routing Policy Schema
Routing policies are defined in zml.toml under the [policy] section, or passed per-call via the routing parameter.
Built-in policies
| Policy | Description |
|---|---|
latency |
Route to the node with the lowest current time-to-first-token estimate |
cost |
Route to the node with the lowest cost_weight that is below queue threshold |
availability |
Route to the node with the most remaining capacity. Maximizes utilization. |
round_robin |
Distribute requests evenly across all healthy nodes regardless of load |
Custom policy example
[policy.custom_cost_latency]
strategy = "weighted"
latency_weight = 0.4
cost_weight = 0.6
queue_threshold = 8 # fallback to latency if queue > 8
Accelerator IDs
Use these type identifiers in your zml.toml accelerator blocks. ZML probes each node at startup and validates the reported hardware against the declared type.
| Type ID | Hardware | Min driver | Notes |
|---|---|---|---|
nvidia_h100 |
NVIDIA H100 SXM / PCIe | 525.85 | CUDA 12.0+ |
nvidia_a100 |
NVIDIA A100 SXM / PCIe | 515.48 | CUDA 11.7+ |
amd_mi300x |
AMD Instinct MI300X | ROCm 6.0 | See cold-start note below |
intel_gaudi2 |
Intel Gaudi2 | 1.12 | Habana Labs driver |
aws_inferentia2 |
AWS Inferentia2 | Neuron SDK 2.18 | inf2.xlarge and above |
Error Codes
ZML raises zml.ZMLError with a code attribute on all non-success responses:
| Code | Meaning |
|---|---|
NO_HEALTHY_NODES |
All registered accelerator nodes failed health checks. Check connectivity and driver status. |
MODEL_NOT_LOADED |
The requested model is not loaded on any available node. Run zml serve for this model first. |
QUEUE_FULL |
All nodes for this model are over the queue threshold. Retry with backoff or scale out nodes. |
AUTH_FAILED |
Invalid or missing API key. Check ZML_API_KEY environment variable. |
TIMEOUT |
Request exceeded the configured timeout value. Increase timeout or reduce prompt length. |
Rate Limits
Rate limits are enforced per API key at the ZML dispatcher level. Default limits for the managed service:
| Plan | Requests/min | Concurrent |
|---|---|---|
| Maker | 60 | 5 |
| Team | 600 | 50 |
| Production | Custom | Custom |
Self-hosted deployments have no rate limits imposed by ZML. Limits can be configured at the reverse proxy layer independently.