🚀 Model Serving & Inference Optimization
Follow a request from queue to first token, size its KV cache, and build a serving plan from measured latency and capacity.
On this page
Before you start
Read Transformers for attention and Quantization for precision and storage. Numerical Computing explains the byte units used below.
By the end, you should be able to calculate cache memory, trace a continuously batched request, distinguish latency from throughput, and identify which measurement would justify a serving optimization.
Start with a request, not a GPU count
A user submits a prompt and waits for an answer. The request passes through authentication, tokenization, a queue, prompt processing, iterative token generation and transport back to the client. A fast model kernel can still sit behind a slow queue or feature service.
Specify the workload first: arrival rate in requests/second, prompt/output token distributions, simultaneous active generations, latency percentiles, model quality and cost budget. Logged-in users, queued requests and active GPU sequences are different counts.
Two phases of generation
Prefill processes prompt tokens and creates the initial cache. Decode produces successive output tokens, appending each token's keys and values. Prefill often has enough parallel work to use matrix hardware well; low-batch decode can be limited by weight or cache traffic. Context length, batch size and architecture can change those bottlenecks.
The KV cache stores attention keys and values that previous tokens already computed. At decode position t, standard full attention still reads earlier keys/values: roughly O(t × d) attention work per head, with head dimension d. Across a generated sequence this attention work remains quadratic in length. The cache avoids redoing the entire prefix's projections and feed-forward computations; it does not make full-attention generation linear.
Work out the cache before choosing concurrency
For conventional MHA/GQA/MQA full attention:
KV bytes = 2 × L × H_kv × d × T × B × b
L = layers, H_kv = key/value heads, d = values per head, T = stored tokens per request, B = equal-length active requests, and b = bytes per stored value. The leading two counts keys and values. For unequal lengths, replace T × B with the sum of stored lengths. This raw-payload model excludes metadata, padding and temporary buffers.
Take L=32, H_kv=8, d=128 and b=2. Each token needs 2 × 32 × 8 × 128 × 2 = 131,072 bytes, or 128 KiB. At 4,096 tokens, one request needs 512 MiB. Sixteen such requests need 8 GiB, in addition to weights and runtime memory.
For a separate 80-layer example with the same KV heads and head dimension, each token costs 320 KiB. At 8,192 tokens that is 2.5 GiB/request, so 100 requests require 250 GiB of raw cache. Parameter count alone cannot tell you this number.
Use binary units consistently: 1 KiB=2¹⁰, 1 MiB=2²⁰, 1 GiB=2³⁰ bytes. Weight counts are often quoted in decimal GB, where 1 GB=10⁹ bytes.
The explorer compares cache layouts under chosen assumptions. MHA, GQA, MQA and MLA are architecture choices with their own training and implementation requirements; changing the comparison does not convert a deployed checkpoint. Real cache allocation also depends on layer attention patterns, sharding and compression metadata.
Allocation: why pages help
Reserving a maximum-length contiguous buffer for every request wastes memory when outputs end early. Paged cache allocation uses fixed-size token blocks and a logical-to-physical block table. A request can grow without moving its whole cache.
With 16-token blocks, a 33-token request occupies three blocks: 48 slots with 15 unused slots at the tail. It stays in those blocks through token 48 and needs a fourth at token 49. Paging reduces reservation and fragmentation waste; it does not reduce the bytes needed for each stored K/V value.
Compatible requests may share an exact token prefix. Prefix reuse must match the model, tokenization, positions, attention/cache configuration and relevant adapter state. Reference counts and copy-on-write protect shared blocks. Cache isolation and eviction policy are part of a multi-user service design.
Scheduling: follow three requests
Static batching fixes membership for a batch; completed requests can return or stream results, but empty slots may wait for the batch to finish. Continuous batching updates membership between iterations.
| Iteration | Active requests | Scheduler decision |
|---|---|---|
| 1 | A and B | Run decode for both |
| 2 | A completes; B continues | Return A's result and release its unshared blocks |
| 3 | B and newly admitted C | Decode B and schedule C's prefill according to the token budget |
An implementation tracks waiting/running requests, token lengths, deadlines, block ownership and cancellation. Before launching work, it reserves required cache capacity and applies a token budget. Chunked prefill can keep a long prompt from monopolizing an iteration, at the cost of more scheduling decisions.
When memory is tight, options include delaying admission, bounded preemption with recomputation, swapping when supported, or rejecting overload explicitly. Silently dropping old cache tokens changes full-attention semantics. Continuous batching improves utilization opportunities, but oversized batches or aggressive prefills can worsen token latency.
Measure latency precisely
- TTFT, time to first token, includes the chosen request-to-first-token boundary: often network, queueing, tokenization, prefill and the first sampling step.
- Inter-token latency, ITL, measures individual gaps between streamed tokens. Report its distribution.
- TPOT, time per output token, is often
(last-token time − first-token time)/(output tokens − 1)for responses with at least two tokens. State the convention. - Throughput is requests/second or tokens/second under a specified workload and latency target. State which tokens are counted.
A response with 100 output tokens, TTFT=0.20 seconds and 99 gaps of 0.04 seconds finishes in 0.20 + 99 × 0.04 = 4.16 seconds. The 40 ms token interval is 25 tokens/second per stream, not 30.
For intuition, streaming 140 GB of weights over an effective 3 TB/s memory path takes 140/3000 ≈ 0.0467 seconds. This is an illustrative bandwidth lower bound if the full payload must be read for that step. Caching, batching, sharding, compression and other work change the actual latency.
Advanced options follow the bottleneck
Speculative decoding lets a cheaper draft propose several tokens, then verifies them in a target-model pass. For target probability p and draft probability q, accept a proposed token with probability min(1,p/q). On rejection, sample from the normalized positive residual max(p−q,0). Under the algorithm's assumptions, this preserves the target sampling distribution, not an identical random output string. Draft cost, acceptance and verification cost decide whether it helps.
Prefill/decode disaggregation uses separate worker pools for the phases. That permits independent scaling but requires cache transfer and coordination. Moving 1 GiB over an effective 25 GB/s path takes at least about 42.95 ms, before protocol and queueing overhead. A short request may not benefit.
Choose TP, PP, quantization, prefix reuse or separate worker pools only after profiling. With TP, check whether KV heads are sharded or replicated by the implementation; total cache divided by GPU count is not always per-device cache. With model replicas, use measured capacity curves and account for warmup and failures.
Check yourself
The 32-layer example has sixteen 4K requests occupying 8 GiB of raw FP16 cache. You switch cache elements to one byte. How much capacity is freed?
Solution: Raw payload falls to 4 GiB, freeing 4 GiB. Scales and packing can reduce the net saving. Model weights do not change, and paging alone would not halve either payload.
Where to go next
Evaluation & Benchmarking helps verify that serving changes preserve useful behavior. Harness Engineering covers reliable orchestration and observable execution around the model.
References
- PagedAttention: block-based cache management and request sharing.
- Speculative decoding: exact target-distribution sampling with draft proposals.
- GQA: grouped-query attention and its cache/quality tradeoff.
- vLLM prefix caching documentation: a concrete implementation and its scope.