How an LLM Predicts the Next Token
Follow a GPT-2-style decoder-only Transformer from input tokens to vectors, repeated blocks, and next-token probabilities. Separate training from inference and see why generation advances one token at a time.
01How an LLM Predicts the Next Token
Concept at a GlanceGPT-2-style structure · components vary by model
An LLM processes input tokens through multiple layers to compute next-token probabilities.
Tokenization produces token IDs. Actual boundaries depend on the tokenizer; this example uses three simplified tokens.
Look up each token embedding and add a position vector, representing token identity and position numerically.
Pass through repeated Transformer blocks. Each block has its own learned weights.
Normalize after the final block. The current last position supplies the vector for choosing the next token.
Output projection and softmax produce next-token probabilities. A high probability does not guarantee a correct answer.
Append the chosen token and start the next computation. Model weights stay fixed during inference.
02 Understand It Simply
For EveryoneAn LLM computes next-token probabilities conditioned on preceding tokens. It appends the chosen token and repeats.
This diagram uses a GPT-2-style architecture with learned positional embeddings and LayerNorm.
Token embeddings plus position vectors pass through repeated Transformer blocks, final normalization, and output projection.
Weights stay fixed during inference.
Token boundaries, dimensions, and layer counts vary by model; this example is simplified.
- –Use it to distinguish model inputs
- –internal computation
- –and outputs
- –or to explain training versus generation
03 Frequently Asked Questions
FAQWhat is How an LLM Predicts the Next Token?+
Follow a GPT-2-style decoder-only Transformer from input tokens to vectors, repeated blocks, and next-token probabilities. Separate training from inference and see why generation advances one token at a time.
Where is How an LLM Predicts the Next Token used?+
Use it to distinguish model inputs, internal computation, and outputs, or to explain training versus generation.
What's a simple analogy for How an LLM Predicts the Next Token?+
An LLM computes next-token probabilities conditioned on preceding tokens. It appends the chosen token and repeats.
