How Probabilities Become Text
Turn output scores into probabilities and choose the next token. Compare how temperature reshapes the distribution, then follow the token-by-token generation loop.
01How Probabilities Become Text
Concept at a GlanceSynthetic values · not real model outputs
Generation repeatedly chooses a token from a next-token distribution. This example uses a synthetic four-token vocabulary.
Project the last position to vocabulary scores. These logits are not probabilities yet.
Softmax turns scores into nonnegative probabilities that sum to one.
A lower temperature concentrates probability on higher scores. Model weights and the ordering of logits stay unchanged.
A higher temperature flattens the distribution. The same draw can select a different token.
Sampling places a draw in cumulative probability intervals. It does not always pick the most probable token.
Append the selected token. A new next-token distribution must be computed for the updated context.
Stop at an end token or the configured length limit. This example stops after two new tokens.
02 Understand It Simply
For EveryoneGeneration repeatedly chooses a next token and appends it to the input. Sampling differs from always selecting the highest-probability token.
Project the last position’s vector to vocabulary logits.
Divide logits by a positive temperature and apply softmax.
A lower temperature concentrates probability but does not guarantee factual accuracy.
This example uses a synthetic four-token vocabulary and fixed draws; it is neither a real model prediction nor a truncated real vocabulary.
Generation stops at an end token or length limit.
- –Use it to understand temperature and sampling settings
- –and set generation length limits
03 Frequently Asked Questions
FAQWhat is How Probabilities Become Text?+
Turn output scores into probabilities and choose the next token. Compare how temperature reshapes the distribution, then follow the token-by-token generation loop.
Where is How Probabilities Become Text used?+
Use it to understand temperature and sampling settings, and set generation length limits.
What's a simple analogy for How Probabilities Become Text?+
Generation repeatedly chooses a next token and appends it to the input. Sampling differs from always selecting the highest-probability token.
