Oh My Algorithm
Concept GuidePrefill · Decode · Memory

Which Computations Does a KV Cache Avoid?

Separate prompt prefill from single-token decode. Reusing earlier K and V avoids recomputation, but each new Q still reads past K and V, and cache storage grows.

01Which Computations Does a KV Cache Avoid?

GPT-2-style structure · components vary by model

A KV cache stores earlier tokens’ K and V in each layer. Compare it with computation without caching.

First process the full prompt to fill K and V. Both paths compute the same prompt positions during prefill.

For a new token, the uncached path recomputes the whole context. The cached path computes Q, K, and V only for the new token.

The new Q attends to both cached and new K. It also combines all V, so the cost of reading past context remains.

Append new K and V to the cache. The next decode step reuses those stored positions.

A conventional full-attention cache grows with context length. Every layer and KV head stores both K and V.

GPT-2 STYLE · DECODER-ONLYWithout cacheK···V···Not computed yet—With KV cacheK···V···Not computed yet—Store K · V per layer and headcache = 2 × layers × KV heads × T × d_headElement count · bytes depend on dtypeSimplified teaching example
1 / 6

In short

A KV cache trades memory for less recomputation. The new Q still requires attention computation and reads past context.

02 Understand It Simply

For Everyone
🔑How It Works

In causal attention, future tokens do not change earlier representations. Each layer can store earlier K and V and reuse them for new tokens.

💡In Plain Words

Prefill builds per-layer K and V for the whole prompt.

Decode computes Q, K, and V for the new token and appends its K and V to the cache.

Its Q still compares with all stored K and combines stored V, so reading past context remains.

A conventional full-attention cache grows with token count.

This diagram shows one simplified layer and head; the cache is distinct from model weights and conversation storage.

📍Where It's Used
  • –Use it to explain long-context memory costs and generation speed
  • –or distinguish prefill and decode bottlenecks

03 Frequently Asked Questions

FAQ
What is Which Computations Does a KV Cache Avoid??+

Separate prompt prefill from single-token decode. Reusing earlier K and V avoids recomputation, but each new Q still reads past K and V, and cache storage grows.

Where is Which Computations Does a KV Cache Avoid? used?+

Use it to explain long-context memory costs and generation speed, or distinguish prefill and decode bottlenecks.

What's a simple analogy for Which Computations Does a KV Cache Avoid??+

In causal attention, future tokens do not change earlier representations. Each layer can store earlier K and V and reuse them for new tokens.