LLM Algorithms
From input tokens to attention, generation, and caching. Learn the 5 topics below step by step with interactive visualizations.
Follow a GPT-2-style decoder-only Transformer from input tokens to vectors, repeated blocks, and next-token probabilities. Separate training from inference and see why generation advances one token at a time.
Decoder-only · InferenceCompute Q–K scores in a small numerical example, mask future positions, normalize the scores, and combine V. Attention weights describe mixing within a layer and head, not the full reason behind a model prediction.
Q · K · VZoom into a GPT-2-style pre-norm block: normalization, attention, residual connections, and MLP. Attention mixes information across tokens; the MLP transforms each position independently.
Pre-norm · Residual · MLPTurn output scores into probabilities and choose the next token. Compare how temperature reshapes the distribution, then follow the token-by-token generation loop.
Softmax · SamplingSeparate prompt prefill from single-token decode. Reusing earlier K and V avoids recomputation, but each new Q still reads past K and V, and cache storage grows.
Prefill · Decode · Memory