Sign In Get Started
Back to Blog

ZML Python SDK 0.3: Async Streaming, Typed Policies, and a New Routing Builder

Version 0.3 of the ZML Python SDK ships async streaming with backpressure handling, a fully typed routing policy builder, and new helpers for accelerator-aware batching. Here is what changed and why.

Python code editor showing ZML SDK 0.3 async streaming interface and typed routing policy builder

ZML Python SDK version 0.3 is available on PyPI now. The headline additions are policy chain configuration, native async/await support, and significantly improved error messages when accelerator routing fails. This post covers the changes, with a focus on what motivated them and what they change for teams already using 0.2.

To install: pip install zml==0.3.0. The 0.3 release introduces no breaking changes to the 0.2 API. Existing inference calls will continue to work unchanged.

Policy Chains: Defining Fallback Behavior in Config

The most requested feature from early-access users was the ability to define accelerator fallback chains without managing the routing logic in application code. In 0.2, if you wanted to prefer H100 and fall back to MI300X, you handled that with try/except around separate infer() calls. That approach worked but it put routing policy in your application code, which makes it hard to change without deploying and creates inconsistency when the same model is served from multiple services.

In 0.3, you define policies in a zml.toml file or inline as Python dictionaries, and reference them by name at call time:

import zml

client = zml.Client(api_key="...")

# Define a policy inline
policy = zml.Policy(
    name="llama3-production",
    accelerator_chain=["H100", "MI300X", "H100_PCIe"],
    latency_target_ms=500,
    cost_priority="balanced"
)

response = client.infer(
    model="meta-llama/llama-3-8b-instruct",
    messages=[{"role": "user", "content": "Hello"}],
    policy=policy
)

The accelerator_chain list defines the preference order. ZML evaluates each accelerator family in order, dispatching to the first that is available and meets the latency target. The cost_priority parameter influences how aggressively ZML will hold to the cheaper option when it is near capacity: "cost" prefers the cheaper option even at some latency risk, "latency" upgrades to the next option before capacity is fully exhausted, "balanced" uses a weighted evaluation.

Policies can also be defined in zml.toml and referenced by name, which is the recommended approach for production configurations where policies should be version-controlled separately from application code:

[policies.llama3-production]
accelerator_chain = ["H100", "MI300X", "H100_PCIe"]
latency_target_ms = 500
cost_priority = "balanced"
response = client.infer(
    model="meta-llama/llama-3-8b-instruct",
    messages=[{"role": "user", "content": "Hello"}],
    policy="llama3-production"  # references zml.toml entry
)

Native Async/Await Support

The 0.2 SDK was synchronous-only. Most ML serving infrastructure today runs async Python applications (FastAPI, aiohttp, async task queues), and calling synchronous ZML methods from an async context required running them in thread pool executors, which added boilerplate and made error handling awkward.

In 0.3, the async path is a first-class citizen. The async client uses the same interface as the sync client, with await:

import asyncio
import zml

async def generate(prompt: str) -> str:
    async with zml.AsyncClient(api_key="...") as client:
        response = await client.infer(
            model="meta-llama/llama-3-8b-instruct",
            messages=[{"role": "user", "content": prompt}],
            policy="llama3-production"
        )
        return response.choices[0].message.content

asyncio.run(generate("Explain tensor parallelism in two sentences"))

Streaming is also supported in the async path via client.infer_stream(), which returns an async generator of token chunks. Cancellation behavior follows Python's standard async cancellation semantics: if you cancel the awaitable, ZML sends a cancellation signal to the serving node and the request is dropped from the queue.

Improved Error Messages

The error messages in 0.2 when routing failed were unhelpful. A common failure looked like this: ZMLRoutingError: Routing failed. There was no indication of which accelerator families were tried, why each was rejected, or what the recommended resolution was.

In 0.3, routing errors include a structured context object that describes what happened at each step of the routing decision:

try:
    response = client.infer(...)
except zml.RoutingError as e:
    print(e.routing_context)
    # {
    #   "policy": "llama3-production",
    #   "evaluated": [
    #     {"family": "H100", "nodes_checked": 3, "rejected_reason": "all_nodes_over_capacity"},
    #     {"family": "MI300X", "nodes_checked": 2, "rejected_reason": "health_score_below_threshold"}
    #   ],
    #   "final_error": "no_viable_node",
    #   "recommendation": "increase fleet capacity or enable queue_on_saturation"
    # }

The routing_context field is also included in the exception's string representation when you log it, so teams using standard Python logging will see the detail without code changes.

Other Changes in 0.3

The response.routing metadata field is now populated for all successful requests, not just when explicitly requested. It includes the selected node ID, accelerator family, policy applied, and queue depth at dispatch time. This is useful for building cost attribution and debugging latency anomalies.

The SDK now validates policy configuration at instantiation time rather than at first call. If a policy references an accelerator family that is not registered in your ZML account, you will get a PolicyConfigError when you create the policy object, not a RoutingError at runtime. This is a significant improvement for catching configuration mistakes before they reach production.

Type stubs (.pyi files) are now included in the package distribution. IDEs with type checking support will now provide accurate autocompletion and type error detection for all public API surfaces.

Migration from 0.2

No code changes are required for existing 0.2 users. The sync Client interface is unchanged. The default behavior when no policy is specified is identical to 0.2. If you were implementing fallback logic manually with try/except, you can migrate that logic to a policy chain at your own pace. The 0.2 pattern continues to work in 0.3.

Full changelog and API reference for 0.3 are available in the docs at zmlai.org/docs/api-reference.html. Bug reports and feature requests go to [email protected] or the community forum.

Start building with the new SDK today

The ZML Python SDK is open source and on PyPI. Install it with pip and follow the quickstart.

Read the quickstart Read the docs