What Changes Inside a Transformer Block?
Zoom into a GPT-2-style pre-norm block: normalization, attention, residual connections, and MLP. Attention mixes information across tokens; the MLP transforms each position independently.
01What Changes Inside a Transformer Block?
Concept at a GlanceGPT-2-style structure · components vary by model
A GPT-2-style block has attention and an MLP, each preceded by normalization and followed by residual addition.
Normalize across each token’s feature dimension. This step does not mix information across tokens.
Attention mixes information from the current and earlier positions. This is the cross-token mixing part of the block.
Add the original input to the attention output. The direct path from input to addition is the residual connection.
Normalize each updated token representation before the MLP. Each block uses separate parameters.
The MLP works independently at each position: expand features, apply GELU, and project back.
Add the MLP output to its residual input. Pass the result onward with token count and feature dimension unchanged.
02 Understand It Simply
For EveryoneAttention mixes information between tokens; the MLP transforms each token’s vector. Residual connections add each sublayer’s input to its output.
A GPT-2-style block computes x + Attention(LayerNorm(x)), then adds MLP(LayerNorm(·)) to that result.
Normalization acts across each token’s feature dimension.
The MLP expands features, applies GELU, and projects back.
This view explains structure without computing real activations.
Other models, such as Llama, use different components including RMSNorm, SwiGLU, and RoPE.
- –Use it to read block diagrams and compare normalization and activation choices across models
03 Frequently Asked Questions
FAQWhat is What Changes Inside a Transformer Block??+
Zoom into a GPT-2-style pre-norm block: normalization, attention, residual connections, and MLP. Attention mixes information across tokens; the MLP transforms each position independently.
Where is What Changes Inside a Transformer Block? used?+
Use it to read block diagrams and compare normalization and activation choices across models.
What's a simple analogy for What Changes Inside a Transformer Block??+
Attention mixes information between tokens; the MLP transforms each token’s vector. Residual connections add each sublayer’s input to its output.
