Sign In Get Started
Back to Blog

Cost Optimization for Production LLM Inference: A Practical Framework

Inference cost is now the biggest line item for teams running LLMs in production. We break down the three levers that actually move the needle: hardware mix, batching strategy, and routing policy.

Infrastructure cost dashboard showing LLM inference optimization metrics across multiple accelerator types

For teams running LLMs in production at any meaningful scale, inference compute has become the largest line item. Not infrastructure management overhead, not training runs, not data pipeline compute: serving real traffic to real users. Understanding where that cost comes from and what actually moves it requires looking at three distinct levers: hardware mix, batching strategy, and routing policy. Each interacts with the others, and optimizing one in isolation often leaves the largest gains on the table.

This post walks through the cost framework we use internally at ZML, with specific numbers where we can give them and honest ranges where we cannot.

Hardware Mix: The Highest-Leverage Decision

The cost of a generated token is driven primarily by which hardware produced it, at what utilization rate, and at what spot vs. on-demand pricing. Hardware mix optimization is the highest-leverage variable because the differences between hardware options are large and the optimization is largely one-time work.

Consider Llama 3 70B serving at moderate load. Running this entirely on H100 SXM on-demand in a major US cloud region costs approximately $2.50 to $3.20 per GPU-hour depending on zone and configuration. Running the same model on MI300X spot in comparable regions runs $1.70 to $2.20 per GPU-hour based on typical spot market rates through 2025. For workloads with low latency sensitivity where batch sizes are above 32, the cost difference compounds significantly at scale.

A concrete example: a production serving deployment processing 40 million tokens per day at typical utilization could see a 20 to 35 percent reduction in hardware cost by routing throughput-oriented traffic to MI300X while reserving H100 SXM for latency-critical paths. The caveat is that this requires MI300X support in your serving stack and a routing layer capable of policy-driven dispatch. Without those, the hardware mix opportunity is theoretical.

The hardware mix lever is also the one that interacts most directly with supply volatility. Optimizing for a single hardware type creates cost exposure when that type moves to on-demand pricing under spot constraints. A multi-hardware routing policy smooths this exposure by maintaining fallback paths at known cost levels.

Batching Strategy: The Ongoing Tuning Work

After hardware selection, continuous batching configuration is the largest controllable cost variable. The core principle: GPU utilization during decode is the primary driver of cost efficiency, and utilization is determined by how well the serving framework fills each forward pass with useful work.

vLLM's PagedAttention and continuous batching handle this reasonably well by default, but the defaults are conservative for most production workloads. The parameters that matter most are max_num_seqs (maximum concurrent sequences per batch), max_num_batched_tokens (total tokens across all sequences in a forward pass), and gpu_memory_utilization (fraction of GPU memory available for KV-cache).

We have found that teams running vLLM with default parameters at moderate load typically achieve 55 to 65 percent GPU utilization. Tuning these parameters for the specific traffic profile of a production workload consistently brings utilization to 75 to 85 percent on the same hardware, reducing cost per token by a proportional amount. The tuning process is iterative: increase max_num_seqs until p99 latency starts creeping, pull back 10 percent, repeat at the next load level.

An important nuance: optimal batching parameters differ between hardware families. H100 SXM and MI300X have different memory bandwidth profiles and different memory capacity. The gpu_memory_utilization setting that maximizes throughput without OOM errors on H100 will not be the same as on MI300X. If you are running a multi-hardware fleet, you need hardware-specific batch configurations, not a single shared setting.

Routing Policy: Where Strategy Meets Execution

Hardware mix and batching parameters define the cost floor for each hardware type. Routing policy determines which traffic actually reaches which hardware and when. A well-designed routing policy combines three dimensions: latency targets, cost priority, and fallback behavior.

A practical policy structure for a mixed H100/MI300X fleet serving both synchronous API calls and background batch jobs might look like this: API calls with latency targets under 500ms route to H100 SXM spots first, with H100 PCIe on-demand as fallback. Batch jobs with multi-second latency budgets route to MI300X spots first, with MI300X on-demand as fallback. If both spot pools are constrained beyond a cost threshold, batch jobs queue rather than escalate to premium hardware.

The queue decision in that last case is worth expanding on. Not all workloads have the same cost tolerance for latency. Background summarization, embedding generation for search indices, and offline analysis jobs are typically tolerant of queuing. Routing these to premium on-demand hardware when spots are unavailable is a policy mistake, not a technical constraint. The policy needs to model both the cost of queuing (delayed output) and the cost of premium hardware (dollars) and make an explicit tradeoff. Without that explicit policy, the default behavior of most serving frameworks is to route to whatever hardware is available at the least resistance, which is usually not cost-optimal.

Metrics That Actually Track Cost

Most teams track GPU utilization as a cost proxy. GPU utilization is necessary but not sufficient. The metrics that more directly track cost efficiency are: cost per million output tokens (CPMT), cost per serving SLA tier (not blended across tiers), and hardware idle time fraction.

CPMT gives you a number that moves when you make optimization changes, regardless of which change caused the movement. It also lets you compare performance across hardware types on equal footing. Tracking cost per SLA tier separately prevents the common issue where batch job cost savings mask latency-tier cost increases. Hardware idle time fraction is the complement of utilization: time the GPU is powered and billed but not processing useful tokens.

We track all three in our ZML dashboard alongside routing decisions. The link between a routing policy change and CPMT movement is visible within a few hours of traffic, which makes iterative optimization practical. Without that link, cost optimization becomes an educated guess exercise.

What the Framework Does Not Change

This framework will not help you if your bottleneck is at a layer above hardware cost. Model size selection, quantization depth, and context length management are architectural decisions that precede hardware optimization. Running Llama 3 70B when Llama 3 8B would meet your accuracy requirements is a cost issue that no hardware mix or routing policy resolves. We are describing optimization within a defined serving configuration, not substituting for that configuration choice.

Similarly, the optimization work described here assumes a stable workload shape. If your traffic distribution changes significantly week over week, batching parameters and routing policies need to be revisited on that cadence, not set once and left. Cost optimization is a process, not a configuration.

Use cost routing to send batch workloads to cheaper nodes automatically

ZML routing policies let you define cost ceilings per request type. The router picks the cheapest node that meets them.

Read routing policy docs Get started