[CHEATSHEET] Understand the overhead of LLM Inference.
Split across prefill and decode:
- Prefill:
- ~60% MLP
- ~35% Attention
- 15% QKV projection
- 10% Attention score & softmax
- 10% Output projection)
- ~5% Layernorm + Other
- Decode:
- ~60β70% Attention
- ~25β30% MLP
- ~5% Other

Prefill phase (compute-bound)
- FFN dominates wall-clock time (since FLOPs β proportional to L).
Decode phase (memory-bound)
- Attention (KV reads/writes) dominates.
Why Different in Decode?
- In decode, each new token attends over all past tokens (L+t).
- KV cache lookups + softmax dominate bandwidth.
- FFN is still FLOP-heavy, but much smaller fraction of total time.
Optimization:
- If youβre optimizing Prefill β focus on MLP kernels (chunking, tensor cores)
- If youβre optimizing Decode β focus on Attention & KV cache (paged attention, compression)