Oh My Algorithm
Concept GuideTokenization · Embedding

From Sentences to Tokens and Vectors

Enter a sentence to inspect real GPT-2 token boundaries and IDs. Select a token to see a four-dimensional teaching example of token and positional embeddings and their sum.

01From Sentences to Tokens and Vectors

Text becomes token IDs, then token and positional vectors are added to form the model input.

Use the GPT-2 tokenizer to split text. Tokens are not necessarily words, and may include leading spaces.

Each token receives its vocabulary ID. Larger IDs do not mean larger meanings or greater similarity.

Look up the embedding row for each ID. Repeated IDs share a vector; the values here are synthetic.

The same token may appear in different positions. GPT-2 separately looks up the vector for each position.

Add token and positional vectors element by element. Select a token in the playground below to inspect the values.

GPT-2 · r50k_baseInput textThe cat satSpaces may belong to tokensTheID 464␠catID 3797␠satID 3332Token embedding-1.751.50.5-0.5E[3797]Position embedding-0.75-0.250.250.75P[1]Model input vector-2.51.250.750.25E[3797] + P[1]4D teaching vectors · not actual model weightsReal token IDs · synthetic vectors
1 / 6

In short

Tokenization maps text to IDs; embedding maps IDs to vectors. Positional representations vary by model.

Try your own sentence

GPT-2 token boundaries and IDs are real. The 4D token and positional vectors below are synthetic teaching values.

Token count: 3

Spaces appear as ␠ and line breaks as ↵. Tokens that do not form complete characters appear as byte fragments.

Selected token: #0 · ID 464

UTF-8: 54 68 65

E[464] · Token embedding

1.750.50-0.75-2.00

P[0] · Position embedding

-1.00-0.75-0.50-0.25

E + P · Model input vector

0.75-0.25-1.25-2.25

The same ID shares a token vector; the same position shares a positional vector. Add the two vectors element by element.

Text decoded from all tokens

The cat sat

02 Understand It Simply

For Everyone
🔑How It Works

Tokenization and embedding are separate steps: tokenization maps text to IDs, while embedding looks up the vector associated with an ID.

💡In Plain Words

This uses GPT-2's byte-level BPE encoding r50k_base.

Spaces may belong to a token, and Korean characters or emoji may span multiple tokens.

Tokens containing incomplete UTF-8 are shown as bytes; decoding the full sequence restores the text.

GPT-2 adds learned positional embeddings to token embeddings.

The four-dimensional vectors here are synthetic teaching values, not GPT-2 weights, meanings, or similarity scores.

Other models use different tokenizers and positional representations.

📍Where It's Used
  • –Compare word and token counts
  • –inspect whitespace and language-dependent boundaries
  • –and distinguish IDs from vectors

03 Frequently Asked Questions

FAQ
What is From Sentences to Tokens and Vectors?+

Enter a sentence to inspect real GPT-2 token boundaries and IDs. Select a token to see a four-dimensional teaching example of token and positional embeddings and their sum.

Where is From Sentences to Tokens and Vectors used?+

Compare word and token counts, inspect whitespace and language-dependent boundaries, and distinguish IDs from vectors.

What's a simple analogy for From Sentences to Tokens and Vectors?+

Tokenization and embedding are separate steps: tokenization maps text to IDs, while embedding looks up the vector associated with an ID.