Oh My Algorithm
Concept GuidePre-norm · Residual · MLP

What Changes Inside a Transformer Block?

Zoom into a GPT-2-style pre-norm block: normalization, attention, residual connections, and MLP. Attention mixes information across tokens; the MLP transforms each position independently.

01What Changes Inside a Transformer Block?

GPT-2-style structure · components vary by model

A GPT-2-style block has attention and an MLP, each preceded by normalization and followed by residual addition.

Normalize across each token’s feature dimension. This step does not mix information across tokens.

Attention mixes information from the current and earlier positions. This is the cross-token mixing part of the block.

Add the original input to the attention output. The direct path from input to addition is the residual connection.

Normalize each updated token representation before the MLP. Each block uses separate parameters.

The MLP works independently at each position: expand features, apply GELU, and project back.

Add the MLP output to its residual input. Pass the result onward with token count and feature dimension unchanged.

GPT-2 STYLE · DECODER-ONLYx · tokens × featuresLayerNormCausal attentioncausal token mixing+ residualLayerNormMLPexpand → GELU → project+ residualPreserve inputAdd through residual pathcausal token mixingcontext mixingtoken → tokenposition-wise MLPInput/output shape preservedSimplified teaching example
1 / 7

In short

Attention mixes across tokens; the MLP transforms each token. Other models use different normalization and activation choices.

02 Understand It Simply

For Everyone
🔑How It Works

Attention mixes information between tokens; the MLP transforms each token’s vector. Residual connections add each sublayer’s input to its output.

💡In Plain Words

A GPT-2-style block computes x + Attention(LayerNorm(x)), then adds MLP(LayerNorm(·)) to that result.

Normalization acts across each token’s feature dimension.

The MLP expands features, applies GELU, and projects back.

This view explains structure without computing real activations.

Other models, such as Llama, use different components including RMSNorm, SwiGLU, and RoPE.

📍Where It's Used
  • –Use it to read block diagrams and compare normalization and activation choices across models

03 Frequently Asked Questions

FAQ
What is What Changes Inside a Transformer Block??+

Zoom into a GPT-2-style pre-norm block: normalization, attention, residual connections, and MLP. Attention mixes information across tokens; the MLP transforms each position independently.

Where is What Changes Inside a Transformer Block? used?+

Use it to read block diagrams and compare normalization and activation choices across models.

What's a simple analogy for What Changes Inside a Transformer Block??+

Attention mixes information between tokens; the MLP transforms each token’s vector. Residual connections add each sublayer’s input to its output.