×

Search anything:

Layerwise timing of LLM

[CHEATSHEET] Understand the overhead of LLM Inference.

Split across prefill and decode:

  • Prefill:
    • ~60% MLP
    • ~35% Attention
      • 15% QKV projection
      • 10% Attention score & softmax
      • 10% Output projection)
    • ~5% Layernorm + Other
  • Decode:
    • ~60–70% Attention
    • ~25–30% MLP
    • ~5% Other

Prefill phase (compute-bound)

  • FFN dominates wall-clock time (since FLOPs β‰ˆ proportional to L).

Decode phase (memory-bound)

  • Attention (KV reads/writes) dominates.

Why Different in Decode?

  • In decode, each new token attends over all past tokens (L+t).
  • KV cache lookups + softmax dominate bandwidth.
  • FFN is still FLOP-heavy, but much smaller fraction of total time.

Optimization:

  • If you’re optimizing Prefill β†’ focus on MLP kernels (chunking, tensor cores)
  • If you’re optimizing Decode β†’ focus on Attention & KV cache (paged attention, compression)
Layerwise timing of LLM
Share this