Foundations
Tokens & Embeddings
FundamentalHow text becomes tokens, tokens become vectors, and attention puts them to work.
What a token actually is
A is the atomic unit a language model reads and writes - not a word, not a character, but a subword chunk produced by a learned compression algorithm called .
Three ways to cut the same sentence
Finer pieces need a smaller vocabulary but make longer sequences; coarser pieces do the reverse. Tap a row. Splits are illustrative.
The middle ground: subwords
Subwords let the model recombine a compact vocabulary into any word it has never seen - here “untokenizable” is rebuilt from pieces it already knows.
Takeaway: subwords sit in the sweet spot - short sequences, a bounded vocabulary, and no unknown words.
How BPE works - one merge at a time
BPE counts every adjacent pair of symbols across its training text, merges the most frequent pair into a new vocabulary entry, and repeats. Real tokenizers start from bytes and run ~50 000 merges; here it runs 12 over a tiny illustrative corpus.
“tokenization”
12 tokens
12 characters = 12 tokens. Watch them fuse.
Most frequent pairs right now
- e + n45
- k + e36
- t + o33
- o + n30
- o + k24
Next merge: e + n. Each count is summed over every word that contains the pair, weighted by how often the word appears.
Training corpus (word × count)
- t·o·k·e·n×11
- t·o·k·e·n·s×8
- t·o·k·e·n·i·z·e×1
- t·o·k·e·n·i·z·a·t·i·o·n×4
- o·r·g·a·n·i·z·a·t·i·o·n×6
- n·a·t·i·o·n×1
- s·t·a·t·i·o·n×2
- t·a·k·e·n×12
- o·f·t·e·n×2
- t·e·n×7
- i·n·t·o×9
- l·i·o·n×9
- i·r·o·n×8
- s·i·z·e×4
- t·i·p×4
Vocabulary
14 chars + 0 merges = 14
- No merges yet - only single characters.
The ordered merge list is the tokenizer: new text is split by replaying these merges in order.
Start: every word is split into single characters. “tokenization” is now t + o + k + e + n + i + z + a + t + i + o + n.
Illustrative corpus and counts, picked so no two pairs tie. Production tokenizers learn their merges from huge text collections, which is why common stems and suffixes end up as single tokens.
Modern tokenizers (GPT-4's cl100k_base, Llama's SentencePiece variants, Gemma's tokenizer) all descend from the BPE idea, with different merge rules, vocabulary sizes, and byte-fallback strategies for rare characters.
See it for yourself
Type any text below and watch a real tokenizer split it into tokens. Switch between four popular tokenizers to see how the same string produces different splits - then hit Compare token counts to see how many tokens each one charges you for. Try prose, code, numbers, emoji, or a non-English sentence.
Tokenizer playground - real tokenizers, in your browser
Pick a tokenizer and type. Each chip is one token, colored by what kind of piece it is, with the integer ID the model actually receives underneath.
BPE · o200k_base · ~200k vocab
Loading GPT-4o tokenizer…
Pieces of one word sit flush against each other; a gap means a new word. These IDs are what the model sees - each one picks a row of its embedding matrix (next section). Hover a chip for its raw form.
Same text, four tokenizers
Each tokenizer's split, fewest tokens first.
Tokenizers run entirely in your browser via Transformers.js; vocabularies are fetched from the Hugging Face Hub on first use and cached. No text leaves your device.
From token to vector
Text tokens are just integers - an ID in a lookup table. The never touches the raw text; it works entirely with dense floating-point vectors. The translation happens in one lookup, after which the vectors flow into attention (see Attention mechanisms for how that works):
From token to vector - one table lookup
Follow “Hello world” from text to IDs to rows of the embedding matrix, and into the model.
1 · Text → IDs
“Hello world”
[9906, 1917]
2 · Embedding matrix
[128,000, 4,096] = ~524 M parameters
3 · Two vectors
row 9906 · Hello
row 1917 · world
values illustrative
4 · Into the model
Transformer
attention layer 1 → …
Tokenize: The tokenizer converts the string into a sequence of integer IDs - [9906, 1917] in one vocabulary, completely different integers in another.
d_model is 4,096 for Llama-3-8B and 8,192 for Llama-3-70B; parameters = vocab_size × d_model at the page's example vocab of 128 000. Rows are drawn to scale, so IDs 1917 and 9906 both land in the first few percent of the table.
The embedding matrix is a large chunk of model parameters: at vocab size 128 000 and d_model 4 096, it accounts for roughly 500 M parameters before you count a single attention layer. This is why large-vocabulary models cost more memory even at the same layer depth. For the full memory accounting, see Model sizing.
See embeddings for real
The lookup above is abstract until you watch it happen. Below, a real embedding model turns your text into actual vectors - first one per token, then a similarity map showing how embeddings place related words close together in space.
Every token becomes a vector
Type a phrase. Each token is mapped to a real 384-dimensional vector by all-MiniLM-L6-v2, drawn as a barcode: one column per dimension, orange positive, blue negative, stronger color = larger magnitude. Stacked strips let you compare tokens column by column.
Loading embedding model (~25 MB, first time only)…
Embeddings capture meaning - the meaning map
A vector on its own is just numbers; the point is distance. Each word below is embedded, then its 384-D vector is projected onto the two directions where these words differ most (PCA) so you can see it. Related words land close together. Pick a word to see its nearest neighbors by cosine similarity (1.0 = same direction, 0 = unrelated) and compare barcodes.
3 to 8 words, comma-separated. Updates as you type.
Embedding words…
This demo uses a single small model (all-MiniLM-L6-v2, 384 dimensions) running entirely in your browser. Every model has its own embedding matrix learned during training - a different model produces different vectors of a different dimension (BERT 768, Llama-3-8B 4 096), so embeddings are not portable across models. The structure you see here - meaning encoded as geometry - is what every model learns, even though the exact numbers differ.
The vocabulary, and why tokenizers differ across models
Every model family ships its own tokenizer trained on its own corpus. The same sentence can become 12 tokens for one model and 19 for another - with real consequences for cost, usage, and quality.
One vocabulary, many consequences
Vocabulary size moves memory between the embedding table and the sequence - and every family's merges differ. Splits shown are illustrative.
Vocabulary size tradeoffs
Smaller vocab (32k-50k): leaner parameter count, but more tokens per sentence and harder coverage of rare languages or technical notation.
Byte fallback: modern tokenizers reserve slots for raw UTF-8 bytes so no character is ever truly unknown - it just costs more tokens.
Why the same text differs
The BPE merge order is determined by the training corpus. A model trained heavily on code will merge programming symbols into fewer tokens than one trained on news.
Whitespace handling varies: some tokenizers treat a leading space as part of a token (“ hello” vs “hello” are different token IDs).
“Hello world” →
You cannot safely share token IDs across model families. A prompt crafted for GPT-4 will have a different ID sequence when fed to Llama or Gemma, even if the text is identical.
Model architecture choices like GQA and MLA affect how the transformer processes those vectors - not how they are tokenized. See Model architectures for that layer of the stack.
Why tokens are the unit of cost and compute
Tokens are not just an internal implementation detail - they are the fundamental unit along which every measurable resource scales: API billing, context window limits, memory for the KV cache, and inference latency.
One token count, four bills
Drag the token count: every resource below scales with it in a straight line. Each cell = 1,000 tokens; layer rows are illustrative.
API billing
×8 the 1,000-token bill
Providers charge per input token and per output token. A 10,000-token context costs ~10× a 1,000-token context - regardless of how many words are in it.
Context window
6.3% of a 128k window
The maximum sequence length (e.g. 128k, 1M) is a token count, not a word or character count. One dense paragraph of code can fill the window faster than the same paragraph of prose.
KV cache memory
×8 the cache, in every layer
Each token in the context occupies a slice of the KV cache for every attention layer. Longer token sequences mean more GPU memory. See the KV cache page for exact formulas.
Throughput & latency
8,000 decode steps · ×8 bandwidth
read as output tokens - one decode step each
Inference systems measure decode speed in tokens per second. A reasoning model that emits 8,000 output tokens per query costs 8× the memory bandwidth of one that emits 1,000.
Characters per token by content type
The rule of thumb for English prose is ~3–4 characters per token, or ~0.75 words per token. But that ratio shifts dramatically by content type. Accurate token estimation matters for cost forecasting and model sizing.
| Content type | Example | Chars / token bar: 0-4 chars | Words / token | Notes |
|---|---|---|---|---|
| Plain English prose | “The model generates a response.” | ~3.5–4 | ~0.75 | Dense but predictable subwords |
| Python source code | def forward(self, x): | ~2.5–3.5 | - | Indentation & operators split frequently |
| JSON / structured data | {"key": "value"} | ~2–3 | - | Punctuation-heavy; brackets each get a token |
| Non-English text (e.g. Chinese) | 模型生成响应 | ~1–2 | - | Many scripts lack subword coverage; each character may be one token |
| Numbers | 3.14159265 | ~1–2 | - | Long integers and decimals split digit by digit |
| URLs / file paths | /api/v2/models/list | ~2–3 | - | Slashes and dots each consume tokens |
Figures are approximate and vary by tokenizer. Always use the model's own tokenizer to get exact counts before committing to a cost estimate or context window design. The KV cache calculator on this site uses token counts as its primary input.
Tokenization gotchas
Tokenization surprises are responsible for a surprising share of production bugs and cost overruns. Here are the most common ones:
Five ways token counts surprise you
Pick a gotcha to see what the tokenizer does. Splits are illustrative.
Leading-space tokens
Most BPE tokenizers merge a leading space into the following word: the token for “ hello” (space+hello) is different from the token for “hello”. This means inserting or removing a space at a boundary changes which IDs are produced - a common source of off-by-one bugs when concatenating prompt fragments programmatically.
Where tokens go next: attention in one picture
Embeddings are where tokenization ends and the model begins. A transformer is a stack of identical blocks - from about 30 to over 100 of them - and each block does two things to every token's vector. First, : each token looks back at the tokens before it and pulls in whatever is relevant, so a word like “it” can pick up the meaning of the noun it refers to. Then a small feed-forward network (an ) refines each vector on its own. After the last block, the newest token's vector is turned into a score for every entry in the vocabulary, and the top choice becomes the next token.
Below is one step of self-attention on toy numbers. The newest token's query is scored against every earlier token's key, turns the scores into weights that add up to 1, and the values are blended by those weights.
Self-attention: score, softmax, blend
The newest token forms a Query, scores it against every token's Key (Q·K / √d), turns the scores into weights with softmax, then mixes their Values by those weights. All numbers are toy 4-d vectors, for illustration only.
| Token | K | Q·K/√d | Softmax weight | V |
|---|---|---|---|---|
Cache cached · frozen | 0.80 | 0.49 score 0.80 | ||
the cached · frozen | −0.41 | 0.14 score −0.41 | ||
keys new · computed now | 0.55 | 0.37 score 0.55 |
Strongest link: "Cache", 2 tokens back - not the neighbor "the". Weights follow what the Q and K vectors encode, not how close the tokens are.
Blend: output = Σ weight × V(bars zoomed in)
The blended vector is what this token carries forward. Only its Query - and its own K and V - were new this step; every other K and V row was read from the cache unchanged.
attn(Q, K, V) = softmax(Q·Kᵀ / √d) · V
A past token's K and V depend only on that token and the weights - never on what comes after. They are frozen the moment the token is processed, so they can be computed once and cached forever. That is precisely what the KV cache stores.
Notice the keys and values marked “cached”: an earlier token's key and value never change, so serving engines keep them instead of recomputing them for every new token. That stored state is the KV cache - the memory bill much of this site is about.
Where to go next
You have followed text all the way into the model: tokens, vectors, and one step of attention. The next pages take the transformer apart - the full block and the architectures built from it, the attention variants that decide memory use, and the KV cache that serving is built around.