Foundations

Tokens & Embeddings

Fundamental

How text becomes tokens, tokens become vectors, and attention puts them to work.

What a token actually is

A is the atomic unit a language model reads and writes - not a word, not a character, but a subword chunk produced by a learned compression algorithm called .

Three ways to cut the same sentence

Finer pieces need a smaller vocabulary but make longer sequences; coarser pieces do the reverse. Tap a row. Splits are illustrative.

The middle ground: subwords

Subwords let the model recombine a compact vocabulary into any word it has never seen - here “untokenizable” is rebuilt from pieces it already knows.

Takeaway: subwords sit in the sweet spot - short sequences, a bounded vocabulary, and no unknown words.

How BPE works - one merge at a time

BPE counts every adjacent pair of symbols across its training text, merges the most frequent pair into a new vocabulary entry, and repeats. Real tokenizers start from bytes and run ~50 000 merges; here it runs 12 over a tiny illustrative corpus.

“tokenization”

12 tokens

12 characters = 12 tokens. Watch them fuse.

Most frequent pairs right now

  • e + n45
  • k + e36
  • t + o33
  • o + n30
  • o + k24

Next merge: e + n. Each count is summed over every word that contains the pair, weighted by how often the word appears.

Training corpus (word × count)

  • t·o·k·e·n×11
  • t·o·k·e·n·s×8
  • t·o·k·e·n·i·z·e×1
  • t·o·k·e·n·i·z·a·t·i·o·n×4
  • o·r·g·a·n·i·z·a·t·i·o·n×6
  • n·a·t·i·o·n×1
  • s·t·a·t·i·o·n×2
  • t·a·k·e·n×12
  • o·f·t·e·n×2
  • t·e·n×7
  • i·n·t·o×9
  • l·i·o·n×9
  • i·r·o·n×8
  • s·i·z·e×4
  • t·i·p×4

Vocabulary

14 chars + 0 merges = 14

  1. No merges yet - only single characters.

The ordered merge list is the tokenizer: new text is split by replaying these merges in order.

1/13Start: every word is split into single characters.

Start: every word is split into single characters. “tokenization” is now t + o + k + e + n + i + z + a + t + i + o + n.

Illustrative corpus and counts, picked so no two pairs tie. Production tokenizers learn their merges from huge text collections, which is why common stems and suffixes end up as single tokens.

Modern tokenizers (GPT-4's cl100k_base, Llama's SentencePiece variants, Gemma's tokenizer) all descend from the BPE idea, with different merge rules, vocabulary sizes, and byte-fallback strategies for rare characters.

See it for yourself

Type any text below and watch a real tokenizer split it into tokens. Switch between four popular tokenizers to see how the same string produces different splits - then hit Compare token counts to see how many tokens each one charges you for. Try prose, code, numbers, emoji, or a non-English sentence.

Tokenizer playground - real tokenizers, in your browser

Pick a tokenizer and type. Each chip is one token, colored by what kind of piece it is, with the integer ID the model actually receives underneath.

BPE · o200k_base · ~200k vocab

Tokens: …Characters: 101Chars / token: -

Loading GPT-4o tokenizer…

word startcontinues a worddigit, punctuation or byte· leading space folded into the token

Pieces of one word sit flush against each other; a gap means a new word. These IDs are what the model sees - each one picks a row of its embedding matrix (next section). Hover a chip for its raw form.

Same text, four tokenizers

Each tokenizer's split, fewest tokens first.

Tokenizers run entirely in your browser via Transformers.js; vocabularies are fetched from the Hugging Face Hub on first use and cached. No text leaves your device.

From token to vector

Text tokens are just integers - an ID in a lookup table. The never touches the raw text; it works entirely with dense floating-point vectors. The translation happens in one lookup, after which the vectors flow into attention (see Attention mechanisms for how that works):

From token to vector - one table lookup

Follow “Hello world” from text to IDs to rows of the embedding matrix, and into the model.

1 · Text → IDs

“Hello world”

Hello→9906
·world→1917

[9906, 1917]

2 · Embedding matrix

d_model = 4,096 columns →128K0rows(vocab)row 9906row 1917

[128,000, 4,096] = ~524 M parameters

3 · Two vectors

row 9906 · Hello

…4,096

row 1917 · world

…4,096

values illustrative

4 · Into the model

vec+pos 0
vec+pos 1

Transformer

attention layer 1 → …

1/4Tokenize: The tokenizer converts the string into a sequence of integer IDs - [9906, 1917] in one vocabulary, completely different integers in another.

Tokenize: The tokenizer converts the string into a sequence of integer IDs - [9906, 1917] in one vocabulary, completely different integers in another.

d_model is 4,096 for Llama-3-8B and 8,192 for Llama-3-70B; parameters = vocab_size × d_model at the page's example vocab of 128 000. Rows are drawn to scale, so IDs 1917 and 9906 both land in the first few percent of the table.

The embedding matrix is a large chunk of model parameters: at vocab size 128 000 and d_model 4 096, it accounts for roughly 500 M parameters before you count a single attention layer. This is why large-vocabulary models cost more memory even at the same layer depth. For the full memory accounting, see Model sizing.

See embeddings for real

The lookup above is abstract until you watch it happen. Below, a real embedding model turns your text into actual vectors - first one per token, then a similarity map showing how embeddings place related words close together in space.

Every token becomes a vector

Type a phrase. Each token is mapped to a real 384-dimensional vector by all-MiniLM-L6-v2, drawn as a barcode: one column per dimension, orange positive, blue negative, stronger color = larger magnitude. Stacked strips let you compare tokens column by column.

Loading embedding model (~25 MB, first time only)…

Embeddings capture meaning - the meaning map

A vector on its own is just numbers; the point is distance. Each word below is embedded, then its 384-D vector is projected onto the two directions where these words differ most (PCA) so you can see it. Related words land close together. Pick a word to see its nearest neighbors by cosine similarity (1.0 = same direction, 0 = unrelated) and compare barcodes.

3 to 8 words, comma-separated. Updates as you type.

Embedding words…

This demo uses a single small model (all-MiniLM-L6-v2, 384 dimensions) running entirely in your browser. Every model has its own embedding matrix learned during training - a different model produces different vectors of a different dimension (BERT 768, Llama-3-8B 4 096), so embeddings are not portable across models. The structure you see here - meaning encoded as geometry - is what every model learns, even though the exact numbers differ.

The vocabulary, and why tokenizers differ across models

Every model family ships its own tokenizer trained on its own corpus. The same sentence can become 12 tokens for one model and 19 for another - with real consequences for cost, usage, and quality.

One vocabulary, many consequences

Vocabulary size moves memory between the embedding table and the sequence - and every family's merges differ. Splits shown are illustrative.

Vocabulary size tradeoffs

Embedding matrix + output softmax~41k rows
“internationalization is hard”6 tokens
internationalization␣is␣hard

Smaller vocab (32k-50k): leaner parameter count, but more tokens per sentence and harder coverage of rare languages or technical notation.

🦙→<0xF0><0x9F><0xA6><0x99>4 byte tokens

Byte fallback: modern tokenizers reserve slots for raw UTF-8 bytes so no character is ever truly unknown - it just costs more tokens.

Why the same text differs

Code-heavy corpus · 6 tokens␣␣␣␣if␣x␣!=␣None:
News-heavy corpus · 10 tokens␣␣␣␣if␣x␣!=␣None:

The BPE merge order is determined by the training corpus. A model trained heavily on code will merge programming symbols into fewer tokens than one trained on news.

␣hello≠hellotwo different IDs

Whitespace handling varies: some tokenizers treat a leading space as part of a token (“ hello” vs “hello” are different token IDs).

“Hello world” →

Llama 399061917
other family??different integers

You cannot safely share token IDs across model families. A prompt crafted for GPT-4 will have a different ID sequence when fed to Llama or Gemma, even if the text is identical.

Model architecture choices like GQA and MLA affect how the transformer processes those vectors - not how they are tokenized. See Model architectures for that layer of the stack.

Why tokens are the unit of cost and compute

Tokens are not just an internal implementation detail - they are the fundamental unit along which every measurable resource scales: API billing, context window limits, memory for the KV cache, and inference latency.

One token count, four bills

Drag the token count: every resource below scales with it in a straight line. Each cell = 1,000 tokens; layer rows are illustrative.

8,000
1,000 (baseline)10,000

API billing

×8 the 1,000-token bill

Providers charge per input token and per output token. A 10,000-token context costs ~10× a 1,000-token context - regardless of how many words are in it.

Context window

6.3% of a 128k window

0128,000 tokens

The maximum sequence length (e.g. 128k, 1M) is a token count, not a word or character count. One dense paragraph of code can fill the window faster than the same paragraph of prose.

KV cache memory

×8 the cache, in every layer

Each token in the context occupies a slice of the KV cache for every attention layer. Longer token sequences mean more GPU memory. See the KV cache page for exact formulas.

Throughput & latency

8,000 decode steps · ×8 bandwidth

read as output tokens - one decode step each

Inference systems measure decode speed in tokens per second. A reasoning model that emits 8,000 output tokens per query costs 8× the memory bandwidth of one that emits 1,000.

Characters per token by content type

The rule of thumb for English prose is ~3–4 characters per token, or ~0.75 words per token. But that ratio shifts dramatically by content type. Accurate token estimation matters for cost forecasting and model sizing.

Content typeExampleChars / token bar: 0-4 charsWords / tokenNotes
Plain English prose“The model generates a response.”~3.5–4~0.75Dense but predictable subwords
Python source codedef forward(self, x):~2.5–3.5-Indentation & operators split frequently
JSON / structured data{"key": "value"}~2–3-Punctuation-heavy; brackets each get a token
Non-English text (e.g. Chinese)模型生成响应~1–2-Many scripts lack subword coverage; each character may be one token
Numbers3.14159265~1–2-Long integers and decimals split digit by digit
URLs / file paths/api/v2/models/list~2–3-Slashes and dots each consume tokens

Figures are approximate and vary by tokenizer. Always use the model's own tokenizer to get exact counts before committing to a cost estimate or context window design. The KV cache calculator on this site uses token counts as its primary input.

Tokenization gotchas

Tokenization surprises are responsible for a surprising share of production bugs and cost overruns. Here are the most common ones:

Five ways token counts surprise you

Pick a gotcha to see what the tokenizer does. Splits are illustrative.

Leading-space tokens

␣hello≠hellodifferent token IDs
"Say" + " hello"Say␣hello
"Say " + "hello"Say␣hello

Most BPE tokenizers merge a leading space into the following word: the token for “ hello” (space+hello) is different from the token for “hello”. This means inserting or removing a space at a boundary changes which IDs are produced - a common source of off-by-one bugs when concatenating prompt fragments programmatically.

Where tokens go next: attention in one picture

Embeddings are where tokenization ends and the model begins. A transformer is a stack of identical blocks - from about 30 to over 100 of them - and each block does two things to every token's vector. First, : each token looks back at the tokens before it and pulls in whatever is relevant, so a word like “it” can pick up the meaning of the noun it refers to. Then a small feed-forward network (an ) refines each vector on its own. After the last block, the newest token's vector is turned into a score for every entry in the vocabulary, and the top choice becomes the next token.

Below is one step of self-attention on toy numbers. The newest token's query is scored against every earlier token's key, turns the scores into weights that add up to 1, and the values are blended by those weights.

Self-attention: score, softmax, blend

The newest token forms a Query, scores it against every token's Key (Q·K / √d), turns the scores into weights with softmax, then mixes their Values by those weights. All numbers are toy 4-d vectors, for illustration only.

1/3New token "keys" attends over 3 tokens
New tokenkeys
Q[1.2, 0.1, −0.3, −0.5]
toy vectors · d = 4
Attention of the new token keys over each token: key, score, softmax weight and value
TokenKSoftmax weightV

Cache

cached · frozen

0.49

score 0.80

the

cached · frozen

0.14

score −0.41

keys

new · computed now

0.37

score 0.55

softmax weights, stackedsum = 1.00
Cache
the
keys

Strongest link: "Cache", 2 tokens back - not the neighbor "the". Weights follow what the Q and K vectors encode, not how close the tokens are.

Blend: output = Σ weight × V(bars zoomed in)

×0.49
+
×0.14
+
×0.37
=
[0.71, 0.17, 0.01, −0.07]

The blended vector is what this token carries forward. Only its Query - and its own K and V - were new this step; every other K and V row was read from the cache unchanged.

attn(Q, K, V) = softmax(Q·Kᵀ / √d) · V

A past token's K and V depend only on that token and the weights - never on what comes after. They are frozen the moment the token is processed, so they can be computed once and cached forever. That is precisely what the KV cache stores.

Notice the keys and values marked “cached”: an earlier token's key and value never change, so serving engines keep them instead of recomputing them for every new token. That stored state is the KV cache - the memory bill much of this site is about.

Where to go next

You have followed text all the way into the model: tokens, vectors, and one step of attention. The next pages take the transformer apart - the full block and the architectures built from it, the attention variants that decide memory use, and the KV cache that serving is built around.