Sign In Get Started
Docs

Quickstart

This guide walks you through installing ZML, setting your API key, writing a minimal routing config, and running your first inference call against your accelerator pool. Most teams finish this in 20 to 30 minutes.

Prerequisites: Python 3.10 or later. At least one supported accelerator node reachable from your environment. See Accelerator IDs for the list of supported targets.

Step 1: Install the SDK

Install the ZML Python SDK from PyPI using pip:

pip install zml

Verify the installation:

python -c "import zml; print(zml.__version__)"
# 0.3.1

Step 2: Set your API key

ZML authenticates using an API key. Set it as an environment variable:

export ZML_API_KEY=your_key_here

You can also pass the key directly when constructing the client, though the environment variable approach is preferred for production deployments:

import zml

client = zml.Client(
    endpoint="https://api.yourinfra.com",
    api_key="your_key_here"
)

Step 3: Write a routing config

Create a zml.toml file in your project root. This tells ZML which accelerator nodes are available and what routing policy to use by default:

[dispatcher]
endpoint = "https://api.yourinfra.com"
default_policy = "latency"

[[accelerators]]
id = "nvidia-h100-node-1"
host = "10.0.1.11"
type = "nvidia_h100"
cost_weight = 1.0

[[accelerators]]
id = "amd-mi300x-node-1"
host = "10.0.1.22"
type = "amd_mi300x"
cost_weight = 0.65
The cost_weight is a relative multiplier used when routing policy is set to "cost". A node with cost_weight = 0.65 is 35 percent cheaper per inference token than a node at 1.0.

Step 4: Run your first inference call

With the config in place, instantiate the client and call infer():

import zml

client = zml.Client()  # reads ZML_API_KEY + zml.toml automatically

result = client.infer(
    model="llama3-70b",
    prompt="What is the capital of France?",
    max_tokens=128
)

print(result.text)
print(f"Served by: {result.metadata.accelerator_id}")

Expected output:

The capital of France is Paris.
Served by: amd-mi300x-node-1

Step 5: Inspect routing metadata

Every response includes a metadata object with routing and performance information:

print(result.metadata.accelerator_id)    # "amd-mi300x-node-1"
print(result.metadata.routing_policy)     # "latency"
print(result.metadata.time_to_first_token) # 0.182  (seconds)
print(result.metadata.queue_depth_at_dispatch) # 2