Which Tokens Does Attention Use?
Compute Q–K scores in a small numerical example, mask future positions, normalize the scores, and combine V. Attention weights describe mixing within a layer and head, not the full reason behind a model prediction.
01Which Tokens Does Attention Use?
Concept at a GlanceSynthetic values · not real model outputs
Attention combines information from visible tokens. Here we use synthetic 2D values for a single head.
Project each token representation into Q, K, and V. Q and K determine mixing weights; V supplies the information to combine.
Take dot products between Q and each K, then divide by the square root of the head dimension.
Set future-position scores to negative infinity. Each token can attend only to itself and earlier positions.
Apply softmax to each row. Future positions have zero weight, and visible weights sum to one.
Combine V using the last row’s weights. The result is a weighted sum, not a copy of a single token.
A real block concatenates and projects multiple head outputs. Heads are not assigned predefined roles.
02 Understand It Simply
For EveryoneAttention compares a token’s Q with visible tokens’ K to determine how their V vectors are combined.
Q, K, and V are learned projections of each layer’s token representations.
Scale QKᵀ by the square root of head dimension, mask future positions with −∞, apply row-wise softmax, then take a weighted sum of V.
Concatenate and project head outputs.
These 2D values are a synthetic, checkable example, not measurements from a real model.
- –Use it to read causal masks and attention matrices
- –and distinguish the roles of Q
- –K
- –and V
03 Frequently Asked Questions
FAQWhat is Which Tokens Does Attention Use??+
Compute Q–K scores in a small numerical example, mask future positions, normalize the scores, and combine V. Attention weights describe mixing within a layer and head, not the full reason behind a model prediction.
Where is Which Tokens Does Attention Use? used?+
Use it to read causal masks and attention matrices, and distinguish the roles of Q, K, and V.
What's a simple analogy for Which Tokens Does Attention Use??+
Attention compares a token’s Q with visible tokens’ K to determine how their V vectors are combined.
