Most of the multi-accelerator serving discussion in the ML community focuses on LLMs. That is understandable given where the majority of inference cost sits in 2026, but it creates a gap for teams running mixed model fleets where vision and text workloads compete for the same hardware pool. The dynamics are different enough that they warrant separate treatment.
Vision models like CLIP, BLIP-2, and Stable Diffusion have different accelerator affinity patterns than LLMs. They have different memory footprints, different sensitivity to batch size, and different precision format behavior. When your hardware pool needs to serve both text and vision workloads, the routing problem is more complex than picking an accelerator family and applying a single policy.
How Vision and LLM Workloads Differ at the Hardware Level
LLM inference is dominated by memory bandwidth: loading large weight tensors and KV-cache entries per forward pass. Performance scales well with more HBM bandwidth and falls off with memory bandwidth bottlenecks. This is why MI300X's large HBM3 pool is attractive for LLM serving, particularly for large models.
Vision model inference has a different profile depending on the model class. CLIP and BLIP-2 embedding generation involves large matrix multiplications over image patch representations, which are compute-bound rather than memory-bound for typical batch sizes. Stable Diffusion inference involves iterative UNet forward passes with attention operations that become memory-bandwidth-sensitive at larger resolutions but compute-sensitive at small resolution and high step counts.
In practice, CLIP embedding generation at batch sizes above 16 shows relatively flat performance differences between H100 and MI300X because both are heavily compute-bound and both exceed the compute requirements of the workload. Stable Diffusion shows more variation: H100's tensor core throughput at FP16 is well-optimized for the UNet attention operations, while MI300X performance depends heavily on which xformers/flash-attention equivalent kernels are available in the ROCm path for your specific diffusion library.
The Queue Competition Problem
The challenge in mixed vision and LLM fleets is that workloads with different hardware affinity often compete for the same node pool when traffic patterns are correlated. If you have a product where image analysis and text generation are triggered together (a document understanding pipeline, for example), demand for vision and LLM inference will spike simultaneously. With a single shared hardware pool and no routing awareness of workload type, the queue competition can cause latency degradation in both workload classes simultaneously.
Consider a concrete scenario: a document processing service uses CLIP to classify document regions and Llama 3 8B to generate summaries. During a traffic spike, both workloads queue for the same nodes. If the routing layer allocates CLIP jobs to H100 SXM nodes that were reserved for latency-sensitive LLM traffic, LLM p95 latency degrades while CLIP completes faster than it would on MI300X. Whether that tradeoff is acceptable depends on which SLA you care more about, and without a routing policy that expresses that preference, the outcome is arbitrary.
Model-Aware Routing Policy Design
ZML routing policies support a model_familyclassification that allows hardware assignment to be differentiated by model type. You can define separate policies for LLM workloads and vision workloads and apply them at the model registration level, so the routing layer makes hardware decisions that reflect the different affinities of each workload class:
import zml
client = zml.Client(api_key="...")
# Vision embedding jobs: cost-primary, high batch tolerance
vision_policy = zml.Policy(
name="clip-embeddings",
accelerator_chain=["MI300X", "H100"],
latency_target_ms=3000,
cost_priority="cost"
)
# LLM serving: latency-primary for interactive use
llm_policy = zml.Policy(
name="llama3-interactive",
accelerator_chain=["H100", "MI300X"],
latency_target_ms=400,
cost_priority="latency"
)
# Dispatch with appropriate policy per model type
clip_response = client.infer(
model="openai/clip-vit-large-patch14",
inputs={"images": image_batch},
policy=vision_policy
)
llm_response = client.infer(
model="meta-llama/llama-3-8b-instruct",
messages=[{"role": "user", "content": prompt}],
policy=llm_policy
)
With separate policies per model class, vision workloads preferentially route to MI300X (cheaper, large memory pool, adequate for compute-bound vision ops), while LLM interactive traffic preferentially routes to H100 for lower decode latency. The hardware pool is shared but the queue competition is managed through policy priority rather than first-come-first-served.
Diffusion Models: The Special Case
Diffusion model serving (Stable Diffusion, SDXL, and newer architectures) is the most hardware-sensitive vision workload and deserves specific attention. The UNet architecture in SD/SDXL has high attention memory bandwidth requirements at 1024x1024 resolution, and the iterative nature of diffusion (typically 20 to 50 forward passes per image) amplifies any per-pass performance difference.
Our internal testing found that SDXL at 1024px, 30 steps, BF16 showed comparable throughput on H100 and MI300X at batch size 4. At batch size 8, MI300X showed slightly lower latency due to memory bandwidth advantage in the UNet attention layers. At batch size 1, H100 was faster by approximately 12 percent, consistent with the latency pattern we see for LLM decoding at low concurrency.
Diffusion model libraries (Diffusers, ComfyUI backend) vary in their ROCm support quality. Diffusers with xformers attention on ROCm is functional but has had kernel coverage gaps in some xformers versions. Teams deploying diffusion models on MI300X should validate their specific library and attention configuration before production rollout.
What Changes When You Rewrite Nothing
The practical goal for a team running both vision and LLM workloads is to avoid writing hardware-specific code paths for each model type. The serving framework configuration should handle hardware differences (batch sizes, precision, memory utilization settings); the routing layer should handle hardware assignment based on model class and traffic priority. Application code should not know which accelerator executed a request.
This is the design constraint we built ZML around. When you register both a CLIP model and a Llama 3 model with separate routing policies, the dispatch to H100 or MI300X is a policy evaluation, not an application code branch. Adding a new accelerator family to the pool requires a policy update, not a code change. When hardware supply changes, you update fallback chains in configuration. The application stays the same.
Teams that have not made this architectural separation will find that every hardware change requires a code review cycle. Teams that have made it will find that hardware changes are operational events, not engineering projects. That is the difference that makes multi-accelerator serving sustainable at the pace GPU supply changes in this industry.
ZML supports vision models across all major accelerator families
Route LLaVA, Qwen-VL, and other vision models across H100, MI300X, and Gaudi2 from a single endpoint.