Inference Optimization Stack / Layered Speedup
1
Continuous Batching
Requests enter/leave batch per token step. Eliminates idle GPU cycles from slow completions.
10-20x
throughput
2
Flash Attention 2
Tiles attention in SRAM. O(n) memory instead of O(n2). 2-4x faster on long sequences.
2-4x
attn speed
3
KV Cache (Paged)
Non-contiguous physical blocks mapped to logical positions. 2-4x more concurrent sequences per memory budget.
2-4x
concurrency
4
Quantization (FP8)
40-50% VRAM reduction at near-identical quality. Enables larger batch sizes on same hardware.
40%
VRAM saved
5
Speculative Decoding
Small draft model proposes tokens; large target verifies in parallel. Add last -- needs stable baseline.
2-3x
token/s
Apply in order 1 to 5
baseline