Which Computations Does a KV Cache Avoid?
Separate prompt prefill from single-token decode. Reusing earlier K and V avoids recomputation, but each new Q still reads past K and V, and cache storage grows.
01Which Computations Does a KV Cache Avoid?
Concept at a GlanceGPT-2-style structure · components vary by model
A KV cache stores earlier tokens’ K and V in each layer. Compare it with computation without caching.
First process the full prompt to fill K and V. Both paths compute the same prompt positions during prefill.
For a new token, the uncached path recomputes the whole context. The cached path computes Q, K, and V only for the new token.
The new Q attends to both cached and new K. It also combines all V, so the cost of reading past context remains.
Append new K and V to the cache. The next decode step reuses those stored positions.
A conventional full-attention cache grows with context length. Every layer and KV head stores both K and V.
02 Understand It Simply
For EveryoneIn causal attention, future tokens do not change earlier representations. Each layer can store earlier K and V and reuse them for new tokens.
Prefill builds per-layer K and V for the whole prompt.
Decode computes Q, K, and V for the new token and appends its K and V to the cache.
Its Q still compares with all stored K and combines stored V, so reading past context remains.
A conventional full-attention cache grows with token count.
This diagram shows one simplified layer and head; the cache is distinct from model weights and conversation storage.
- –Use it to explain long-context memory costs and generation speed
- –or distinguish prefill and decode bottlenecks
03 Frequently Asked Questions
FAQWhat is Which Computations Does a KV Cache Avoid??+
Separate prompt prefill from single-token decode. Reusing earlier K and V avoids recomputation, but each new Q still reads past K and V, and cache storage grows.
Where is Which Computations Does a KV Cache Avoid? used?+
Use it to explain long-context memory costs and generation speed, or distinguish prefill and decode bottlenecks.
What's a simple analogy for Which Computations Does a KV Cache Avoid??+
In causal attention, future tokens do not change earlier representations. Each layer can store earlier K and V and reuse them for new tokens.
