From Sentences to Tokens and Vectors
Enter a sentence to inspect real GPT-2 token boundaries and IDs. Select a token to see a four-dimensional teaching example of token and positional embeddings and their sum.
01From Sentences to Tokens and Vectors
Concept at a GlanceText becomes token IDs, then token and positional vectors are added to form the model input.
Use the GPT-2 tokenizer to split text. Tokens are not necessarily words, and may include leading spaces.
Each token receives its vocabulary ID. Larger IDs do not mean larger meanings or greater similarity.
Look up the embedding row for each ID. Repeated IDs share a vector; the values here are synthetic.
The same token may appear in different positions. GPT-2 separately looks up the vector for each position.
Add token and positional vectors element by element. Select a token in the playground below to inspect the values.
Try your own sentence
GPT-2 token boundaries and IDs are real. The 4D token and positional vectors below are synthetic teaching values.
Token count: 3
Spaces appear as ␠ and line breaks as ↵. Tokens that do not form complete characters appear as byte fragments.
Selected token: #0 · ID 464
UTF-8: 54 68 65
E[464] · Token embedding
P[0] · Position embedding
E + P · Model input vector
The same ID shares a token vector; the same position shares a positional vector. Add the two vectors element by element.
Text decoded from all tokens
The cat sat
02 Understand It Simply
For EveryoneTokenization and embedding are separate steps: tokenization maps text to IDs, while embedding looks up the vector associated with an ID.
This uses GPT-2's byte-level BPE encoding r50k_base.
Spaces may belong to a token, and Korean characters or emoji may span multiple tokens.
Tokens containing incomplete UTF-8 are shown as bytes; decoding the full sequence restores the text.
GPT-2 adds learned positional embeddings to token embeddings.
The four-dimensional vectors here are synthetic teaching values, not GPT-2 weights, meanings, or similarity scores.
Other models use different tokenizers and positional representations.
- –Compare word and token counts
- –inspect whitespace and language-dependent boundaries
- –and distinguish IDs from vectors
03 Frequently Asked Questions
FAQWhat is From Sentences to Tokens and Vectors?+
Enter a sentence to inspect real GPT-2 token boundaries and IDs. Select a token to see a four-dimensional teaching example of token and positional embeddings and their sum.
Where is From Sentences to Tokens and Vectors used?+
Compare word and token counts, inspect whitespace and language-dependent boundaries, and distinguish IDs from vectors.
What's a simple analogy for From Sentences to Tokens and Vectors?+
Tokenization and embedding are separate steps: tokenization maps text to IDs, while embedding looks up the vector associated with an ID.
