Oh My Algorithm
Concept GuideQ · K · V

Which Tokens Does Attention Use?

Compute Q–K scores in a small numerical example, mask future positions, normalize the scores, and combine V. Attention weights describe mixing within a layer and head, not the full reason behind a model prediction.

01Which Tokens Does Attention Use?

Synthetic values · not real model outputs

Attention combines information from visible tokens. Here we use synthetic 2D values for a single head.

Project each token representation into Q, K, and V. Q and K determine mixing weights; V supplies the information to combine.

Take dot products between Q and each K, then divide by the square root of the head dimension.

Set future-position scores to negative infinity. Each token can attend only to itself and earlier positions.

Apply softmax to each row. Future positions have zero weight, and visible weights sum to one.

Combine V using the last row’s weights. The result is a weighted sum, not a copy of a single token.

A real block concatenates and projects multiple head outputs. Heads are not assigned predefined roles.

GPT-2 STYLE · DECODER-ONLYThecatsatSynthetic Q · K · VQ100111K100111V100221causal attentionTheThecatcatsatsat·········Rows = Q · columns = KSimplified teaching example
1 / 7

In short

Attention weights describe mixing in this layer and head; they are not a complete explanation of a prediction.

02 Understand It Simply

For Everyone
🔑How It Works

Attention compares a token’s Q with visible tokens’ K to determine how their V vectors are combined.

💡In Plain Words

Q, K, and V are learned projections of each layer’s token representations.

Scale QKᵀ by the square root of head dimension, mask future positions with −∞, apply row-wise softmax, then take a weighted sum of V.

Concatenate and project head outputs.

These 2D values are a synthetic, checkable example, not measurements from a real model.

📍Where It's Used
  • –Use it to read causal masks and attention matrices
  • –and distinguish the roles of Q
  • –K
  • –and V

03 Frequently Asked Questions

FAQ
What is Which Tokens Does Attention Use??+

Compute Q–K scores in a small numerical example, mask future positions, normalize the scores, and combine V. Attention weights describe mixing within a layer and head, not the full reason behind a model prediction.

Where is Which Tokens Does Attention Use? used?+

Use it to read causal masks and attention matrices, and distinguish the roles of Q, K, and V.

What's a simple analogy for Which Tokens Does Attention Use??+

Attention compares a token’s Q with visible tokens’ K to determine how their V vectors are combined.