Oh My Algorithm
Concept GuideSoftmax · Sampling

How Probabilities Become Text

Turn output scores into probabilities and choose the next token. Compare how temperature reshapes the distribution, then follow the token-by-token generation loop.

01How Probabilities Become Text

Synthetic values · not real model outputs

Generation repeatedly chooses a token from a next-token distribution. This example uses a synthetic four-token vocabulary.

Project the last position to vocabulary scores. These logits are not probabilities yet.

Softmax turns scores into nonnegative probabilities that sum to one.

A lower temperature concentrates probability on higher scores. Model weights and the ordering of logits stay unchanged.

A higher temperature flattens the distribution. The same draw can select a different token.

Sampling places a draw in cumulative probability intervals. It does not always pick the most probable token.

Append the selected token. A new next-token distribution must be computed for the updated context.

Stop at an end token or the configured length limit. This example stops after two new tokens.

GPT-2 STYLE · DECODER-ONLYCurrent inputThecatsatTransformer → logitsNext-token scoreson2by1near0<eos>-1T = 1 · softmax(logits / T)Simplified teaching example
1 / 8

In short

Temperature reshapes the distribution; it does not guarantee factual accuracy or quality.

02 Understand It Simply

For Everyone
🔑How It Works

Generation repeatedly chooses a next token and appends it to the input. Sampling differs from always selecting the highest-probability token.

💡In Plain Words

Project the last position’s vector to vocabulary logits.

Divide logits by a positive temperature and apply softmax.

A lower temperature concentrates probability but does not guarantee factual accuracy.

This example uses a synthetic four-token vocabulary and fixed draws; it is neither a real model prediction nor a truncated real vocabulary.

Generation stops at an end token or length limit.

📍Where It's Used
  • –Use it to understand temperature and sampling settings
  • –and set generation length limits

03 Frequently Asked Questions

FAQ
What is How Probabilities Become Text?+

Turn output scores into probabilities and choose the next token. Compare how temperature reshapes the distribution, then follow the token-by-token generation loop.

Where is How Probabilities Become Text used?+

Use it to understand temperature and sampling settings, and set generation length limits.

What's a simple analogy for How Probabilities Become Text?+

Generation repeatedly chooses a next token and appends it to the input. Sampling differs from always selecting the highest-probability token.