Quickstart
This guide walks you through installing ZML, setting your API key, writing a minimal routing config, and running your first inference call against your accelerator pool. Most teams finish this in 20 to 30 minutes.
Step 1: Install the SDK
Install the ZML Python SDK from PyPI using pip:
pip install zml
Verify the installation:
python -c "import zml; print(zml.__version__)"
# 0.3.1
Step 2: Set your API key
ZML authenticates using an API key. Set it as an environment variable:
export ZML_API_KEY=your_key_here
You can also pass the key directly when constructing the client, though the environment variable approach is preferred for production deployments:
import zml
client = zml.Client(
endpoint="https://api.yourinfra.com",
api_key="your_key_here"
)
Step 3: Write a routing config
Create a zml.toml file in your project root. This tells ZML which accelerator nodes are available and what routing policy to use by default:
[dispatcher]
endpoint = "https://api.yourinfra.com"
default_policy = "latency"
[[accelerators]]
id = "nvidia-h100-node-1"
host = "10.0.1.11"
type = "nvidia_h100"
cost_weight = 1.0
[[accelerators]]
id = "amd-mi300x-node-1"
host = "10.0.1.22"
type = "amd_mi300x"
cost_weight = 0.65
cost_weight is a relative multiplier used when routing policy is set to "cost". A node with cost_weight = 0.65 is 35 percent cheaper per inference token than a node at 1.0.
Step 4: Run your first inference call
With the config in place, instantiate the client and call infer():
import zml
client = zml.Client() # reads ZML_API_KEY + zml.toml automatically
result = client.infer(
model="llama3-70b",
prompt="What is the capital of France?",
max_tokens=128
)
print(result.text)
print(f"Served by: {result.metadata.accelerator_id}")
Expected output:
The capital of France is Paris.
Served by: amd-mi300x-node-1
Step 5: Inspect routing metadata
Every response includes a metadata object with routing and performance information:
print(result.metadata.accelerator_id) # "amd-mi300x-node-1"
print(result.metadata.routing_policy) # "latency"
print(result.metadata.time_to_first_token) # 0.182 (seconds)
print(result.metadata.queue_depth_at_dispatch) # 2