Foundations
Model Architectures
IntermediateFrom the decoder-only transformer through MoE, state-space models, and multimodal - and what each choice costs at inference time.
Prerequisite: this page references attention, KV cache, and head dimensions throughout. If those are new, start with attention in one picture on Tokens & Embeddings; the attention variants themselves are on Attention Mechanisms.
The decoder-only transformer: today's default
Almost every production LLM you encounter today is a decoder-only - the architecture family that started with GPT and now covers Llama, Mistral, Gemma, Qwen, and hundreds of open and proprietary variants. It won out over encoder-decoder (BERT, T5) and pure-encoder designs for generative tasks because autoregressive generation maps cleanly onto it: each new token attends to all prior tokens and produces the next one.
A typical model stacks 32–96 identical transformer blocks. Each block has four logical components:
Inside one transformer block
The full path from text to next-token probabilities, with tensor shapes on every box. Step a token through, or tap any box.
Input text
1/16Raw prompt text.
Text & stream
Input text
Raw prompt text. Nothing numeric yet.
Legend - Llama 3.1 8B
- B
- Batch size - sequences the server runs togethervaries
- L
- Sequence length - tokens so far (+1 per generated token)varies
- D
- Model width, d_model (hidden size)4,096
- H
- Query heads / KV heads (grouped-query)32 / 8, head size 128
- F
- Feed-forward width14,336
- V
- Vocabulary size128,256
- N
- Blocks (layers)32
Values from Meta's config.json on Hugging Face.
Each block reads from the residual stream and adds back at the ⊕. That additive structure is why deep stacks train stably and why individual layers can be quantized or pruned without collapse.
The “” is the key abstraction: a vector of dimension d_model (e.g., 4096 for Llama-3 8B, 8192 for 70B) that persists per token through all layers. Every attention head and every reads from and writes back to it. This additive structure lets deep networks train reliably and is why you can prune or quantize individual layers without catastrophic collapse.
See Attention Mechanisms for a deep dive into multi-head, grouped-query, and multi-latent attention variants.
Dense vs sparse: capacity vs compute
The single most important distinction in modern model design is how many parameters fire per token. Dense models activate everything; sparse models route each token through a fraction of their capacity. The tradeoff changes the memory footprint, the compute cost, and the serving strategy.
Dense vs sparse - what sits in memory vs what runs
Bar length = parameters resident in GPU memory (to scale, in billions). Filled = the weights that actually compute for the current token. Step through a sentence and watch the MoE row light different experts.
Dense, 13B-class · every weight fires, e.g. Llama-3, Mistral 7B
memory ~13B · runs ~13B (100%)
Sparse MoE - Mixtral 8x7B · Mixtral, DeepSeek-V3, Qwen3-MoE, Llama 4 · first thin block (S) = shared weights
memory ~47B · runs ~13B (28%)
Dense, 47B-class · same memory as Mixtral, all of it runs
memory ~47B · runs ~47B (100%)
"The" → router picks E1 + E7: ~13B of ~47B weights run
Mixtral gets the FLOPs of a ~13B model with the memory bill of a ~47B one.
Any token can pick any expert, so every expert must stay resident. Sparsity cuts compute per token, not the memory you have to provision.
Active share per token, MoEs on this page
- Mixtral 8x7B28% (13/47B)
- Qwen3 235B-A22B9.4% (22/235B)
- DeepSeek-V35.5% (37/671B)
- Llama 4 Maverick4.3% (17/400B)
- Kimi K23.2% (32/1,000B)
- Any dense model100%
Computed from the approximate total / active parameter counts in the table below. The shared vs per-expert split in the Mixtral bar is derived from ~47B total and ~13B active with 8 experts, top-2 - illustrative.
The critical implication for infrastructure: a sparse model's device memory requirement is set by its total parameters, not its active ones. Mixtral 8x7B has ~47 B total parameters and must reside in memory as such - even though only ~13 B of those weights execute per token. You get the FLOPs of a 13 B model with the memory bill of a 47 B one.
Mixture-of-Experts (MoE)
replaces each dense MLP layer with a set of expert sub-networks plus a lightweight router. The router decides which experts handle each token; the rest sit idle. This lets a model scale capacity (total parameters) far beyond what its compute budget (active parameters) would normally allow.
Routing a token: router → top-2 experts → weighted sum
The router scores every expert for each token. Only the top 2 run, and their outputs are blended by the renormalized scores. Anchored on Mixtral 8x7B (8 experts, top-2).
Expert load so far
1 token · 2 picks
Green = picked by the current token; dashed line = perfectly even share. A load-balancing term keeps routing spread out - without it the router can collapse onto a few experts.
Parameters - Mixtral 8x7B
Compute follows the active parameters; memory holds all of them, since any token may pick any expert.
"The" → E1 (0.56) + E7 (0.44)
Four mechanics behind every MoE layer
Pick one to see it in a small picture. Scores and token counts are illustrative.
Router scores, one token
out = 0.58·E3(x) + 0.42·E5(x)
Two experts run and their outputs are blended by the renormalized scores - some redundancy. The other 6 experts stay idle for this token.
Top-k gating
A lightweight linear router produces a score for each expert. The top-k experts (typically k=1 or k=2) are selected and their outputs are summed with learned weights. k=1 minimizes compute; k=2 adds redundancy.
Real-world MoE models
Total vs active parameters, on one scale
The whole bar must be resident in memory; only the green part runs for each token. Tap a model for its routing.
DeepSeek-V3 / R1
~671 B total~37 B activeRouting: 256 routed + shared, top-8
Fine-grained experts + Multi-head Latent Attention (MLA) shrink the KV cache to ~70 KB/token. Auxiliary-loss-free balancing and FP8 training make it the efficiency reference point - see the callout below.
Why is DeepSeek-V3 so much more efficient?
671 B total · 37 B activeDeepSeek-V3 isn't just “another MoE.” It stacks four optimizations that compound - which is why it matches frontier dense models at a fraction of the training and serving cost.
KV cache per token, same scale
Multi-head Latent Attention (MLA)
Instead of caching full Keys and Values per head, MLA compresses them into a small shared low-rank latent vector and reconstructs on the fly. The KV cache drops to ~70 KB/token - roughly 4–7× smaller than a comparable GQA model - so the same GPU serves far longer contexts.
Busy experts get a lower bias, idle ones a higher one - no extra loss term (illustrative loads)
Auxiliary-loss-free load balancing
Most MoEs add a load-balancing penalty that quietly fights the main objective and hurts quality. DeepSeek-V3 balances experts with a dynamic per-expert bias term instead - no auxiliary loss, so capacity is spread evenly without taxing accuracy.
256 routed - top-8 fire
Fine-grained shared + routed experts
256 small routed experts (top-8) plus always-on shared experts. Smaller experts specialize more precisely, and the shared experts capture common patterns so routed ones don’t waste capacity relearning them - more effective capacity per active parameter.
FP8 training + Multi-Token Prediction
The first frontier model trained largely in 8-bit floating point, roughly halving memory and bandwidth versus BF16. Multi-Token Prediction trains the model to predict several tokens ahead, improving sample efficiency and enabling faster speculative decoding at inference.
The headline number: compresses the KV cache to ~70 KB per token versus ~327–516 KB for comparable models - a 4–7× reduction that lets one GPU serve far longer contexts. The memory chart below makes this concrete.
Parameter counts are approximate; architectures continue to evolve. See Model Sizing & Parallelism for the memory arithmetic.
State-space models: Mamba and beyond
Attention's quadratic cost - O(n²) compute and O(n) KV cache memory with sequence length - is a fundamental constraint. replace the attention mechanism with a learned linear recurrence: a compact hidden state that is updated token-by-token without ever materialising a full attention matrix.
Attention (Transformer)
O(n²) computeEvery new token attends to all previous tokens. The Keys & Values of past tokens are stored in the KV cache, which grows linearly with context length.
State-space model (Mamba)
O(1) memoryEach token updates one fixed-size hidden state and is discarded. No KV cache - memory is constant no matter how long the context grows, but the state must compress all history (weaker exact recall).
From recurrence to Mamba-2
Three steps in how SSMs evolved, and why they matter at long context. Tap a stage.
Linear recurrence
SSMs propagate a fixed-size hidden state forward in time. Each new token updates the state in O(1) - inference is a recurrence, not a full attention over history.
How an SSM keeps its memory
An SSM doesn't store the history - it keeps a running, compressed summary of it in one fixed-size state vector h. Each token updates that state with a linear recurrence, then is discarded:
One state, updated token by token
Step through a sentence and watch each token fold into the fixed-size state h. Values are illustrative.
ht = A · ht-1 + B · xt // fold the new token into the state
yt = C · ht // read an output out of the state
How much of each token is still in h
Retention Ā
Step size Δ
State h after token 1
6
numbers in h - at every token
1
tokens a KV cache would hold now
The tradeoff: the transformer keeps an exact transcript (lossless recall, but it grows forever); the SSM keeps fixed-size notes (constant memory, but it can't quote an arbitrary old token verbatim). That weakness is why frontier models go hybrid.
KV cache consumption, mathematically
The contrast is exact, not hand-wavy: the transformer's KV cache is linear in context length, the SSM's state is constant, and a hybrid pays only for its few attention layers. Slide the context and watch the gap open.
Memory vs context - Transformer vs SSM vs Hybrid
The transformer caches K & V for every token at every layer: 2·L·H·d·b·n, a straight line in context n. An SSM keeps one fixed state per layer (L·d_inner·d_state·b) that ignores n. Drag the cursor along the chart.
Context 128K tokens
Transformer · KV cache · grows with n
32.0 GB
256 KB/token × 128K
Hybrid · 8/64 attention layers
4.01 GB
32 KB/token × 128K + 14.0 MB
SSM · fixed state · flat
16.0 MB
16.0 MB at any n
At 128K context the transformer needs ~8.0× the hybrid and ~2,048× the SSM - and the SSM line never moves.
Only n moves the transformer term, so the gap widens without bound. In this config the transformer's KV cache alone fills an 80 GB H100 at about 320K tokens, before any weights are loaded. That linear-vs-flat split is why long-context systems lean on SSM and hybrid designs.
| Context | Transformer | Hybrid | SSM |
|---|---|---|---|
| 1K | 256 MB | 46.0 MB | 16.0 MB |
| 2K | 512 MB | 78.0 MB | 16.0 MB |
| 4K | 1.00 GB | 142 MB | 16.0 MB |
| 8K | 2.00 GB | 270 MB | 16.0 MB |
| 16K | 4.00 GB | 526 MB | 16.0 MB |
| 32K | 8.00 GB | 1.01 GB | 16.0 MB |
| 64K | 16.0 GB | 2.01 GB | 16.0 MB |
| 128K | 32.0 GB | 4.01 GB | 16.0 MB |
| 256K | 64.0 GB | 8.01 GB | 16.0 MB |
| 512K | 128.0 GB | 16.0 GB | 16.0 MB |
| 1M | 256.0 GB | 32.0 GB | 16.0 MB |
Illustrative ~7-70B-class config (BF16): L=64, KV heads=8, d_head=128, d_inner=8192, d_state=16; the hybrid caches K/V in 8 of 64 layers. Values are exact for this config and per single request (they scale with batch). Real models vary; try them in the KV tiering calculator.
Hybrid models (attention + linear / SSM layers)
Pure SSMs sacrifice some quality on tasks that benefit from exact long-range recall. Hybrids interleave full-attention layers with fixed-state layers - -2 in NVIDIA's Nemotron 3 and IBM's Granite 4.0, in Qwen3.5, and in Kimi K3 and GLM-5.3-Flash - getting most of the quality from the attention layers and most of the memory efficiency from the rest. Only the attention layers produce KV cache entries; the others keep a fixed-size state per sequence.
Hybrid stacks - where the memory goes
Each strip is one model's layer stack. Attention layers write K/V for every token, so their cache grows with context. Linear layers (Gated DeltaNet, KDA, Mamba-2) keep one fixed-size state per sequence. Pick a context length and compare each hybrid with the same model if every layer were attention.
Qwen3.5-397B-A17B
Alibaba Qwen · Feb 2026
45 Gated DeltaNet : 15 full attention
state 180 MiB + KV 3.75 GiB = 3.93 GiB
all-attention 15 GiB - 3.8× less than all-attention
Kimi K3
Moonshot AI · Jul 2026
69 KDA : 24 MLA
state 414 MiB + KV 3.38 GiB = 3.78 GiB
all-attention 13.1 GiB - 3.5× less than all-attention
GLM-5.3-Flash
Z.ai · Aug 2026
34 KDA : 11 MLA
state 136 MiB + KV ~1.38 GiB = 1.51 GiB
all-attention 5.63 GiB - 3.7× less than all-attention
Nemotron 3 Super
NVIDIA
40 Mamba-2 : 8 GQA
state 160 MiB + KV 1 GiB = 1.16 GiB
all-attention 6 GiB - 5.2× less than all-attention
Nemotron 3 Ultra
NVIDIA
48 Mamba-2 : 12 GQA
state 384 MiB + KV 1.5 GiB = 1.88 GiB
all-attention 7.5 GiB - 4.0× less than all-attention
IBM Granite 4.0-H
IBM · Oct 2025
9 Mamba-2 : 1 attention
Jamba
AI21 Labs · 2024
7 Mamba : 1 attention
One sequence; bf16 KV, fp32 state where the config says so. Bars share one scale per context length. At 8K the fixed state can outweigh the KV it replaces - the savings arrive at long context. Ratios are exact; attention positions in the strips are spread evenly for drawing (Qwen3.5 repeats 3 : 1 exactly). Granite 4.0-H and Jamba are shown by ratio only.
SSMs eliminate the KV cache and replace it with a fixed-size recurrent state. See KV Cache for why that matters when context grows long.
Vision & multimodal: encoders bolted onto LLMs
Modern don't retrain the LLM from scratch with images. The dominant pattern: take a pretrained image encoder, project its outputs into the LLM's token space, and concatenate them with the text sequence. The LLM's weights are largely frozen; only the projector (and optionally the encoder) is fine-tuned. The LLM then “reads” image patches as if they were text tokens.
From pixels to tokens - how a VLM reads an image
Step through the pipeline and watch one image turn into hundreds of tokens in the LLM's context.
336 / 14 = 24 × 24 = 576 patches
01 · Split into patches
The image is cut into a grid of fixed-size patches (14 px here). Each patch will become one token.
LLM context sequence
not in the LLM yet
Split into patches: The image is cut into a grid of fixed-size patches (14 px here). Each patch will become one token.
Patch counts follow from resolution ÷ 14 px patch size: 224 px gives 256 tokens and 336 px gives 576, the typical range for one image. Video at 1 fps for 10 seconds is 10 frames, easily 2,500+ tokens. The 7-token text prompt is illustrative; dynamic-resolution models change the patch count per image.
Representative VLMs
| Model | Modality (+ text) | Vision encoder | Notes |
|---|---|---|---|
| Qwen3-VL | ImageVideo | ViT, dynamic resolution + mRoPE | Alibaba's current flagship VLM (2025). Dynamic resolution adjusts patch count to image size, and time-aware position encoding handles video - strong agentic and long-context vision. |
| Llama 4 Maverick | Image | Native multimodal (MoE backbone) | Meta (2025). Vision trained into a 400B/17B-active MoE rather than bolted on, with a 128K+ context for high-resolution images. |
| Gemini 2.5 Pro / Flash | ImageVideoAudio | Undisclosed (native) | Google DeepMind (2025). End-to-end multimodal with very long native context; video understanding is a primary use case. |
| InternVL3 | ImageVideo | InternViT + Qwen2.5 backbone | OpenGVLab (2025). Native multimodal pretraining (not adapter-bolted) for strong cross-modal grounding; a leading open-weight VLM. |
| GPT-4o | ImageVideoAudio | Undisclosed (native) | OpenAI. Audio and vision trained end-to-end; popularized real-time multimodal interaction. |
| LLaVA-1.5 | Image | CLIP ViT-L | The reference open VLM that introduced the connector/projector pattern most open models still follow. |
How many tokens is an image, a video, a recording?
Current models cut images into 14-16 px patches, then merge each 2×2 block into one token - so one token covers a 28-32 px square, and the count grows with the image's area. Others cut the image into fixed tiles plus a thumbnail, or squeeze every image into a fixed budget. The same picture can cost ten times more tokens on one model than another.
| Tokens for one image | 1024 × 1024 | 1920 × 1080 |
|---|---|---|
| Gemma 3fixed 896 px input | 256 | 256 |
| Gemini 3fixed budget, default setting | 1,120 | 1,120 |
| Qwen3-VL16 px patches, 2×2 merge | 1,024 | 2,040 |
| OpenAI GPT-5.x32 px patches | 1,229 | 2,448 |
| Claude (standard)28 px per token, capped | 1,369 | 1,560 |
| Qwen2.5-VL, GLM-4.5V14 px patches, 2×2 merge | 1,369 | 2,691 |
| Llama 4336 px tiles + a global view | ~2,467 | ~2,322 |
| InternVL3448 px tiles + a thumbnail | 2,560 | 2,304 |
Video is a stack of frames sampled at a fixed rate, and audio arrives at a steady token rate. Both grow with running time, which is where the numbers get large:
| Tokens for | 1 minute | 10 minutes | 1 hour |
|---|---|---|---|
| Audio (speech)25 tokens per second | 1.5K | 15K | 90K |
| Video, Gemini low resolution1 frame/s, 66 tokens a frame + audio | 5.9K | 59K | 353K |
| Video, Gemini default1 frame/s, 258 tokens a frame + audio | 17.4K | 174K | 1.04M |
| Video, Qwen3-VL at 2 fps720p, two frames merged per token grid | 43K | 432K | 2.6M |
| Video, Qwen3-VL default capsat most 768 frames per video | 43K | ~106K | ~106K |
Why video blows up memory
Image, video and audio tokens sit in the exactly like text, at the same bytes per token. Qwen3-VL-235B caches about 188 KB per token in BF16. One minute of video at 2 frames per second is 43K tokens - about 8 GB of KV. An hour at Gemini's 1 frame per second is about a million tokens, or roughly 200 GB of KV for a single session - more than two H100s hold in total. A hundred users each sharing one minute of video need about 830 GB of KV on top of 470 GB of weights.
Sample fewer, smaller frames
Dropping the frame rate or switching to a low-resolution mode cuts tokens about 4x. Qwen3-VL caps a video at 768 frames by default, so an hour costs about 106K tokens instead of 2.6M.
Prune or merge visual tokens
Neighboring patches and frames repeat each other. FastV drops about half the visual tokens after the second layer, and token merging (ToMe) folds similar tokens together, with small accuracy losses in benchmarks.
Pick a hybrid-attention model
Models that keep full attention in only some layers cache far less per token - Qwen3.5-397B keeps a full KV cache in 15 of its 60 layers, about 6x less per token than Qwen3-VL-235B.
Tier the KV cache
A long video session that sits idle is a prime candidate for offloading to host memory or shared storage, then reloading when the user returns - see Context Memory.
Serving adds one more stage: the vision encoder runs once per image or frame during prefill. Engines cache its output by image hash so a repeated image is not re-encoded, and vLLM, SGLang and NVIDIA Dynamo can run the encoder on separate GPUs from prefill and decode (encode-prefill-decode disaggregation). What passes between them is small - about 17 MB of embeddings for a 1080p screenshot - while the KV cache it turns into is about 0.4 GB.
Size it: open the KV planner with 100 users each sharing 10 minutes of video on Qwen3-VL-235B, then change the resolution, frame rate or clip length.
Architecture drives the infrastructure bill
Every architecture choice above maps directly to a serving cost. The views below summarize the key dimensions: how memory scales, how sequence-length scaling behaves, and what that means when you size a cluster.
The numbers: weights vs KV cache, by model
Two separate memory bills dominate serving. Weights are a fixed, one-time cost (loaded once, shared across all requests). KV cache is a per-request, per-token cost that grows with context length and concurrency - and it is where architecture choices show up most dramatically.
Two memory bills: weights vs KV cache
Both bills share one fixed scale from 16 GB to 700 TB. The weights never change; add concurrent 128K-token requests and watch the KV cache grow until it dwarfs them.
8 B params · KV/token ~128 KB
47 B / 13 B active params · KV/token ~128 KB
70 B params · KV/token ~320 KB
405 B params · KV/token ~516 KB
671 B / 37 B active params · KV/token ~70 KB
varies params · KV/token no KV cache
Weights: scales with params
No KV cache - a constant state at any concurrency
At one request, weights dominate every row. KV cache is per request, though - add users to see it take over.
Read across the bottom three rows: a dense 405 B model needs ~63 GB of KV cache for a single 128K-token request, while DeepSeek-V3 - a far larger 671 B model - needs only ~9 GB thanks to MLA, and an SSM needs essentially none. KV cache, not weights, is what limits how many long-context users you can pack onto a GPU. Weights assume BF16 (2 bytes/param); KV-cache figures are per single request and scale linearly with batch size.
Qualitative comparison
Five families, three questions
How cost grows with sequence length, how much of the model runs per token, and what that means for sizing. Curves show shape only.
Dense Transformer
Llama-3 70B, Mistral 7B
Memory proportional to full param count. Predictable but expensive at scale. KV cache grows linearly with batch × context.
Sparse MoE
Mixtral 8x7B, DeepSeek-V3
Must load all params onto device(s). Needs expert parallelism. FLOPs/token low but memory footprint is the total - not the active - parameter count.
SSM / Mamba
Mamba-2, Falcon Mamba
No KV cache; fixed-size hidden state instead. Memory flat w.r.t. context length - transformative for very long sequences.
Hybrid (attention + linear / SSM)
Qwen3.5, Kimi K3, Nemotron 3, Jamba
KV cache only for the full-attention layers (1 in 4 to 1 in 10), plus a fixed-size state per sequence. Much less memory at long context, but the state is harder to prefix-cache and offload.
Multimodal VLM
Qwen3-VL, Llama 4 Maverick, InternVL3
Image tokens occupy sequence budget. High-res inputs produce 500–2000 extra tokens per image - KV cache and prefill cost spike accordingly.
Gauge: solid = share of weights active per token; lighter tail = the range across models; striped = depends on the backbone.
Parameter counts and active-param ratios are illustrative; exact figures vary by configuration and version. Confirm against model cards and provider documentation before sizing production clusters.
Dig deeper
Architecture determines memory layout, serving topology, and cost - explore these pages to put the numbers on the table.