Foundations

Model Architectures

Intermediate

From the decoder-only transformer through MoE, state-space models, and multimodal - and what each choice costs at inference time.

Prerequisite: this page references attention, KV cache, and head dimensions throughout. If those are new, start with attention in one picture on Tokens & Embeddings; the attention variants themselves are on Attention Mechanisms.

The decoder-only transformer: today's default

Almost every production LLM you encounter today is a decoder-only - the architecture family that started with GPT and now covers Llama, Mistral, Gemma, Qwen, and hundreds of open and proprietary variants. It won out over encoder-decoder (BERT, T5) and pure-encoder designs for generative tasks because autoregressive generation maps cleanly onto it: each new token attends to all prior tokens and produces the next one.

A typical model stacks 32–96 identical transformer blocks. Each block has four logical components:

Inside one transformer block

The full path from text to next-token probabilities, with tensor shapes on every box. Step a token through, or tap any box.

Example model

Input text

1/16

Raw prompt text.

Decoder block× 32K, VInput text“The cat sat on the”Tokenizer(B, L)BPE · 128,256-token vocabularyToken embedding(B, L, 4,096)looks up 1 row of a V × D tableRMSNorm(B, L, 4,096)pre-norm: rescales a copy of the streamMasked self-attention(B, L, 4,096)grouped-query: 32 Q heads share 8 KV heads ·RoPE on Q, K(B, L, 4,096)RMSNorm(B, L, 4,096)rescales a copy before the feed-forwardFeed-forward (SwiGLU)(B, L, 4,096)gate & up D → F, SiLU gate, down → D · dense(B, L, 4,096)Final RMSNorm(B, L, 4,096)once, after the last blockLM head (linear)(B, L, 128,256)D → V scores (logits)Softmax(B, L, 128,256)scores → probabilities that sum to 1Next-token probabilitiesmat0.41floor0.17sofa0.08illustrativeKV cache128 KiB per token
1/16Text arrives: “The cat sat on the”.

Text & stream

Input text

Raw prompt text. Nothing numeric yet.

Legend - Llama 3.1 8B

B
Batch size - sequences the server runs togethervaries
L
Sequence length - tokens so far (+1 per generated token)varies
D
Model width, d_model (hidden size)4,096
H
Query heads / KV heads (grouped-query)32 / 8, head size 128
F
Feed-forward width14,336
V
Vocabulary size128,256
N
Blocks (layers)32

Values from Meta's config.json on Hugging Face.

Tokens & embeddingsAttentionFeed-forwardNorm & residualOutputresidual add

Each block reads from the residual stream and adds back at the ⊕. That additive structure is why deep stacks train stably and why individual layers can be quantized or pruned without collapse.

The “” is the key abstraction: a vector of dimension d_model (e.g., 4096 for Llama-3 8B, 8192 for 70B) that persists per token through all layers. Every attention head and every reads from and writes back to it. This additive structure lets deep networks train reliably and is why you can prune or quantize individual layers without catastrophic collapse.

See Attention Mechanisms for a deep dive into multi-head, grouped-query, and multi-latent attention variants.

Dense vs sparse: capacity vs compute

The single most important distinction in modern model design is how many parameters fire per token. Dense models activate everything; sparse models route each token through a fraction of their capacity. The tradeoff changes the memory footprint, the compute cost, and the serving strategy.

Dense vs sparse - what sits in memory vs what runs

Bar length = parameters resident in GPU memory (to scale, in billions). Filled = the weights that actually compute for the current token. Step through a sentence and watch the MoE row light different experts.

Dense, 13B-class · every weight fires, e.g. Llama-3, Mistral 7B

memory ~13B · runs ~13B (100%)

all 13B

Sparse MoE - Mixtral 8x7B · Mixtral, DeepSeek-V3, Qwen3-MoE, Llama 4 · first thin block (S) = shared weights

memory ~47B · runs ~13B (28%)

E1
E2
E3
E4
E5
E6
E7
E8

Dense, 47B-class · same memory as Mixtral, all of it runs

memory ~47B · runs ~47B (100%)

all 47B
1/5"The" → router picks E1 + E7: ~13B of ~47B weights run

"The" → router picks E1 + E7: ~13B of ~47B weights run

runs for this token (compute)loaded but idle (memory only)

Mixtral gets the FLOPs of a ~13B model with the memory bill of a ~47B one.

Any token can pick any expert, so every expert must stay resident. Sparsity cuts compute per token, not the memory you have to provision.

Active share per token, MoEs on this page

  • Mixtral 8x7B28%
  • Qwen3 235B-A22B9.4%
  • DeepSeek-V35.5%
  • Llama 4 Maverick4.3%
  • Kimi K23.2%
  • Any dense model100%

Computed from the approximate total / active parameter counts in the table below. The shared vs per-expert split in the Mixtral bar is derived from ~47B total and ~13B active with 8 experts, top-2 - illustrative.

The critical implication for infrastructure: a sparse model's device memory requirement is set by its total parameters, not its active ones. Mixtral 8x7B has ~47 B total parameters and must reside in memory as such - even though only ~13 B of those weights execute per token. You get the FLOPs of a 13 B model with the memory bill of a 47 B one.

Mixture-of-Experts (MoE)

replaces each dense MLP layer with a set of expert sub-networks plus a lightweight router. The router decides which experts handle each token; the rest sit idle. This lets a model scale capacity (total parameters) far beyond what its compute budget (active parameters) would normally allow.

Routing a token: router → top-2 experts → weighted sum

The router scores every expert for each token. Only the top 2 run, and their outputs are blended by the renormalized scores. Anchored on Mixtral 8x7B (8 experts, top-2).

TheRoutergate scores (illustrative)0.380.300.560.44E1E2E3E4E5E6E7E8+weighted sumout = 0.56·E1(x) + 0.44·E7(x)

Expert load so far

1 token · 2 picks

E1E2E3E4E5E6E7E811

Green = picked by the current token; dashed line = perfectly even share. A load-balancing term keeps routing spread out - without it the router can collapse onto a few experts.

Parameters - Mixtral 8x7B

~13B active per token~47B resident in memory

Compute follows the active parameters; memory holds all of them, since any token may pick any expert.

1/8"The" → E1 (0.56) + E7 (0.44)

"The" → E1 (0.56) + E7 (0.44)

Four mechanics behind every MoE layer

Pick one to see it in a small picture. Scores and token counts are illustrative.

Router scores, one token

.05
.08
.31
.06
.22
.12
.09
.07
E1E2E3E4E5E6E7E8

out = 0.58·E3(x) + 0.42·E5(x)

Two experts run and their outputs are blended by the renormalized scores - some redundancy. The other 6 experts stay idle for this token.

Top-k gating

A lightweight linear router produces a score for each expert. The top-k experts (typically k=1 or k=2) are selected and their outputs are summed with learned weights. k=1 minimizes compute; k=2 adds redundancy.

Real-world MoE models

Total vs active parameters, on one scale

The whole bar must be resident in memory; only the green part runs for each token. Tap a model for its routing.

0250 B500 B750 B1 T

DeepSeek-V3 / R1

~671 B total~37 B active

Routing: 256 routed + shared, top-8

Fine-grained experts + Multi-head Latent Attention (MLA) shrink the KV cache to ~70 KB/token. Auxiliary-loss-free balancing and FP8 training make it the efficiency reference point - see the callout below.

Why is DeepSeek-V3 so much more efficient?

671 B total · 37 B active

DeepSeek-V3 isn't just “another MoE.” It stacks four optimizations that compound - which is why it matches frontier dense models at a fraction of the training and serving cost.

Multi-head Latent Attention (MLA)

Instead of caching full Keys and Values per head, MLA compresses them into a small shared low-rank latent vector and reconstructs on the fly. The KV cache drops to ~70 KB/token - roughly 4–7× smaller than a comparable GQA model - so the same GPU serves far longer contexts.

−b
+b
−b
+b

Busy experts get a lower bias, idle ones a higher one - no extra loss term (illustrative loads)

Auxiliary-loss-free load balancing

Most MoEs add a load-balancing penalty that quietly fights the main objective and hurts quality. DeepSeek-V3 balances experts with a dynamic per-expert bias term instead - no auxiliary loss, so capacity is spread evenly without taxing accuracy.

Fine-grained shared + routed experts

256 small routed experts (top-8) plus always-on shared experts. Smaller experts specialize more precisely, and the shared experts capture common patterns so routed ones don’t waste capacity relearning them - more effective capacity per active parameter.

FP8 training + Multi-Token Prediction

The first frontier model trained largely in 8-bit floating point, roughly halving memory and bandwidth versus BF16. Multi-Token Prediction trains the model to predict several tokens ahead, improving sample efficiency and enabling faster speculative decoding at inference.

The headline number: compresses the KV cache to ~70 KB per token versus ~327–516 KB for comparable models - a 4–7× reduction that lets one GPU serve far longer contexts. The memory chart below makes this concrete.

Parameter counts are approximate; architectures continue to evolve. See Model Sizing & Parallelism for the memory arithmetic.

State-space models: Mamba and beyond

Attention's quadratic cost - O(n²) compute and O(n) KV cache memory with sequence length - is a fundamental constraint. replace the attention mechanism with a learned linear recurrence: a compact hidden state that is updated token-by-token without ever materialising a full attention matrix.

Attention (Transformer)

O(n²) compute
Thecatsatonmat

Every new token attends to all previous tokens. The Keys & Values of past tokens are stored in the KV cache, which grows linearly with context length.

State-space model (Mamba)

O(1) memory
fixedstateThecatsatonmat

Each token updates one fixed-size hidden state and is discarded. No KV cache - memory is constant no matter how long the context grows, but the state must compress all history (weaker exact recall).

From recurrence to Mamba-2

Three steps in how SSMs evolved, and why they matter at long context. Tap a stage.

Linear recurrence

SSMs propagate a fixed-size hidden state forward in time. Each new token updates the state in O(1) - inference is a recurrence, not a full attention over history.

How an SSM keeps its memory

An SSM doesn't store the history - it keeps a running, compressed summary of it in one fixed-size state vector h. Each token updates that state with a linear recurrence, then is discarded:

One state, updated token by token

Step through a sentence and watch each token fold into the fixed-size state h. Values are illustrative.

ht = A · ht-1 + B · xt // fold the new token into the state

yt = C · ht // read an output out of the state

How much of each token is still in h

Retention Ā

Step size Δ

State h after token 1

6

numbers in h - at every token

1

tokens a KV cache would hold now

1/8Token 1 “The” folded in with Δ 0.20; older tokens fade by Ā = 0.9 per step.

The tradeoff: the transformer keeps an exact transcript (lossless recall, but it grows forever); the SSM keeps fixed-size notes (constant memory, but it can't quote an arbitrary old token verbatim). That weakness is why frontier models go hybrid.

KV cache consumption, mathematically

The contrast is exact, not hand-wavy: the transformer's KV cache is linear in context length, the SSM's state is constant, and a hybrid pays only for its few attention layers. Slide the context and watch the gap open.

Memory vs context - Transformer vs SSM vs Hybrid

The transformer caches K & V for every token at every layer: 2·L·H·d·b·n, a straight line in context n. An SSM keeps one fixed state per layer (L·d_inner·d_state·b) that ignores n. Drag the cursor along the chart.

Context 128K tokens

064 GB128 GB192 GB256 GB0256K512K768K1Mcontext length n (tokens)memory per request80 GB - one H100 (KV alone)TransformerHybridSSM32.0 GB4.01 GB16.0 MB

Transformer · KV cache · grows with n

32.0 GB

256 KB/token × 128K

Hybrid · 8/64 attention layers

4.01 GB

32 KB/token × 128K + 14.0 MB

SSM · fixed state · flat

16.0 MB

16.0 MB at any n

At 128K context the transformer needs ~8.0× the hybrid and ~2,048× the SSM - and the SSM line never moves.

Only n moves the transformer term, so the gap widens without bound. In this config the transformer's KV cache alone fills an 80 GB H100 at about 320K tokens, before any weights are loaded. That linear-vs-flat split is why long-context systems lean on SSM and hybrid designs.

Memory per request by context length
ContextTransformerHybridSSM
1K256 MB46.0 MB16.0 MB
2K512 MB78.0 MB16.0 MB
4K1.00 GB142 MB16.0 MB
8K2.00 GB270 MB16.0 MB
16K4.00 GB526 MB16.0 MB
32K8.00 GB1.01 GB16.0 MB
64K16.0 GB2.01 GB16.0 MB
128K32.0 GB4.01 GB16.0 MB
256K64.0 GB8.01 GB16.0 MB
512K128.0 GB16.0 GB16.0 MB
1M256.0 GB32.0 GB16.0 MB

Illustrative ~7-70B-class config (BF16): L=64, KV heads=8, d_head=128, d_inner=8192, d_state=16; the hybrid caches K/V in 8 of 64 layers. Values are exact for this config and per single request (they scale with batch). Real models vary; try them in the KV tiering calculator.

Hybrid models (attention + linear / SSM layers)

Pure SSMs sacrifice some quality on tasks that benefit from exact long-range recall. Hybrids interleave full-attention layers with fixed-state layers - -2 in NVIDIA's Nemotron 3 and IBM's Granite 4.0, in Qwen3.5, and in Kimi K3 and GLM-5.3-Flash - getting most of the quality from the attention layers and most of the memory efficiency from the rest. Only the attention layers produce KV cache entries; the others keep a fixed-size state per sequence.

Hybrid stacks - where the memory goes

Each strip is one model's layer stack. Attention layers write K/V for every token, so their cache grows with context. Linear layers (Gated DeltaNet, KDA, Mamba-2) keep one fixed-size state per sequence. Pick a context length and compare each hybrid with the same model if every layer were attention.

Context per sequence
  • Qwen3.5-397B-A17B

    Alibaba Qwen · Feb 2026

    45 Gated DeltaNet : 15 full attention

    state 180 MiB + KV 3.75 GiB = 3.93 GiB

    all-attention 15 GiB - 3.8× less than all-attention

  • Kimi K3

    Moonshot AI · Jul 2026

    69 KDA : 24 MLA

    state 414 MiB + KV 3.38 GiB = 3.78 GiB

    all-attention 13.1 GiB - 3.5× less than all-attention

  • GLM-5.3-Flash

    Z.ai · Aug 2026

    34 KDA : 11 MLA

    state 136 MiB + KV ~1.38 GiB = 1.51 GiB

    all-attention 5.63 GiB - 3.7× less than all-attention

  • Nemotron 3 Super

    NVIDIA

    40 Mamba-2 : 8 GQA

    state 160 MiB + KV 1 GiB = 1.16 GiB

    all-attention 6 GiB - 5.2× less than all-attention

  • Nemotron 3 Ultra

    NVIDIA

    48 Mamba-2 : 12 GQA

    state 384 MiB + KV 1.5 GiB = 1.88 GiB

    all-attention 7.5 GiB - 4.0× less than all-attention

IBM Granite 4.0-H

IBM · Oct 2025

9 Mamba-2 : 1 attention

Jamba

AI21 Labs · 2024

7 Mamba : 1 attention

attention layer / its KV cache (grows with context)linear layer / its fixed state (flat)same model, attention in every layer

One sequence; bf16 KV, fp32 state where the config says so. Bars share one scale per context length. At 8K the fixed state can outweigh the KV it replaces - the savings arrive at long context. Ratios are exact; attention positions in the strips are spread evenly for drawing (Qwen3.5 repeats 3 : 1 exactly). Granite 4.0-H and Jamba are shown by ratio only.

SSMs eliminate the KV cache and replace it with a fixed-size recurrent state. See KV Cache for why that matters when context grows long.

Vision & multimodal: encoders bolted onto LLMs

Modern don't retrain the LLM from scratch with images. The dominant pattern: take a pretrained image encoder, project its outputs into the LLM's token space, and concatenate them with the text sequence. The LLM's weights are largely frozen; only the projector (and optionally the encoder) is fine-tuned. The LLM then “reads” image patches as if they were text tokens.

From pixels to tokens - how a VLM reads an image

Step through the pipeline and watch one image turn into hundreds of tokens in the LLM's context.

336 / 14 = 24 × 24 = 576 patches

01 · Split into patches

The image is cut into a grid of fixed-size patches (14 px here). Each patch will become one token.

LLM context sequence

not in the LLM yet

1/5Split into patches

Split into patches: The image is cut into a grid of fixed-size patches (14 px here). Each patch will become one token.

Patch counts follow from resolution ÷ 14 px patch size: 224 px gives 256 tokens and 336 px gives 576, the typical range for one image. Video at 1 fps for 10 seconds is 10 frames, easily 2,500+ tokens. The 7-token text prompt is illustrative; dynamic-resolution models change the patch count per image.

Representative VLMs

ModelModality (+ text)Vision encoderNotes
Qwen3-VL
ImageVideo
ViT, dynamic resolution + mRoPEAlibaba's current flagship VLM (2025). Dynamic resolution adjusts patch count to image size, and time-aware position encoding handles video - strong agentic and long-context vision.
Llama 4 Maverick
Image
Native multimodal (MoE backbone)Meta (2025). Vision trained into a 400B/17B-active MoE rather than bolted on, with a 128K+ context for high-resolution images.
Gemini 2.5 Pro / Flash
ImageVideoAudio
Undisclosed (native)Google DeepMind (2025). End-to-end multimodal with very long native context; video understanding is a primary use case.
InternVL3
ImageVideo
InternViT + Qwen2.5 backboneOpenGVLab (2025). Native multimodal pretraining (not adapter-bolted) for strong cross-modal grounding; a leading open-weight VLM.
GPT-4o
ImageVideoAudio
Undisclosed (native)OpenAI. Audio and vision trained end-to-end; popularized real-time multimodal interaction.
LLaVA-1.5
Image
CLIP ViT-LThe reference open VLM that introduced the connector/projector pattern most open models still follow.

How many tokens is an image, a video, a recording?

Current models cut images into 14-16 px patches, then merge each 2×2 block into one token - so one token covers a 28-32 px square, and the count grows with the image's area. Others cut the image into fixed tiles plus a thumbnail, or squeeze every image into a fixed budget. The same picture can cost ten times more tokens on one model than another.

Tokens for one image1024 × 10241920 × 1080
Gemma 3fixed 896 px input256256
Gemini 3fixed budget, default setting1,1201,120
Qwen3-VL16 px patches, 2×2 merge1,0242,040
OpenAI GPT-5.x32 px patches1,2292,448
Claude (standard)28 px per token, capped1,3691,560
Qwen2.5-VL, GLM-4.5V14 px patches, 2×2 merge1,3692,691
Llama 4336 px tiles + a global view~2,467~2,322
InternVL3448 px tiles + a thumbnail2,5602,304

Video is a stack of frames sampled at a fixed rate, and audio arrives at a steady token rate. Both grow with running time, which is where the numbers get large:

Tokens for1 minute10 minutes1 hour
Audio (speech)25 tokens per second1.5K15K90K
Video, Gemini low resolution1 frame/s, 66 tokens a frame + audio5.9K59K353K
Video, Gemini default1 frame/s, 258 tokens a frame + audio17.4K174K1.04M
Video, Qwen3-VL at 2 fps720p, two frames merged per token grid43K432K2.6M
Video, Qwen3-VL default capsat most 768 frames per video43K~106K~106K

Why video blows up memory

Image, video and audio tokens sit in the exactly like text, at the same bytes per token. Qwen3-VL-235B caches about 188 KB per token in BF16. One minute of video at 2 frames per second is 43K tokens - about 8 GB of KV. An hour at Gemini's 1 frame per second is about a million tokens, or roughly 200 GB of KV for a single session - more than two H100s hold in total. A hundred users each sharing one minute of video need about 830 GB of KV on top of 470 GB of weights.

Sample fewer, smaller frames

Dropping the frame rate or switching to a low-resolution mode cuts tokens about 4x. Qwen3-VL caps a video at 768 frames by default, so an hour costs about 106K tokens instead of 2.6M.

Prune or merge visual tokens

Neighboring patches and frames repeat each other. FastV drops about half the visual tokens after the second layer, and token merging (ToMe) folds similar tokens together, with small accuracy losses in benchmarks.

Pick a hybrid-attention model

Models that keep full attention in only some layers cache far less per token - Qwen3.5-397B keeps a full KV cache in 15 of its 60 layers, about 6x less per token than Qwen3-VL-235B.

Tier the KV cache

A long video session that sits idle is a prime candidate for offloading to host memory or shared storage, then reloading when the user returns - see Context Memory.

Serving adds one more stage: the vision encoder runs once per image or frame during prefill. Engines cache its output by image hash so a repeated image is not re-encoded, and vLLM, SGLang and NVIDIA Dynamo can run the encoder on separate GPUs from prefill and decode (encode-prefill-decode disaggregation). What passes between them is small - about 17 MB of embeddings for a 1080p screenshot - while the KV cache it turns into is about 0.4 GB.

Size it: open the KV planner with 100 users each sharing 10 minutes of video on Qwen3-VL-235B, then change the resolution, frame rate or clip length.

Architecture drives the infrastructure bill

Every architecture choice above maps directly to a serving cost. The views below summarize the key dimensions: how memory scales, how sequence-length scaling behaves, and what that means when you size a cluster.

The numbers: weights vs KV cache, by model

Two separate memory bills dominate serving. Weights are a fixed, one-time cost (loaded once, shared across all requests). KV cache is a per-request, per-token cost that grows with context length and concurrency - and it is where architecture choices show up most dramatically.

Two memory bills: weights vs KV cache

Both bills share one fixed scale from 16 GB to 700 TB. The weights never change; add concurrent 128K-token requests and watch the KV cache grow until it dwarfs them.

Concurrent 128K requests
weights (BF16) KV cache
Llama 3.1 8BDense · GQA

8 B params · KV/token ~128 KB

16 GB
~16 GB1.0× the weights
Mixtral 8x7BMoE · GQA

47 B / 13 B active params · KV/token ~128 KB

94 GB
~16 GB0.17× the weights
Llama 3.1 70BDense · GQA

70 B params · KV/token ~320 KB

140 GB
~40 GB0.29× the weights
Llama 3.1 405BDense · GQA

405 B params · KV/token ~516 KB

810 GB
~63 GB0.08× the weights
DeepSeek-V3MoE · MLA

671 B / 37 B active params · KV/token ~70 KB

1.34 TB (≈671 GB FP8)
~9 GB0.01× the weights
Mamba-2 (SSM)State-space

varies params · KV/token no KV cache

Weights: scales with params

No KV cache - a constant state at any concurrency

At one request, weights dominate every row. KV cache is per request, though - add users to see it take over.

Read across the bottom three rows: a dense 405 B model needs ~63 GB of KV cache for a single 128K-token request, while DeepSeek-V3 - a far larger 671 B model - needs only ~9 GB thanks to MLA, and an SSM needs essentially none. KV cache, not weights, is what limits how many long-context users you can pack onto a GPU. Weights assume BF16 (2 bytes/param); KV-cache figures are per single request and scale linearly with batch size.

Qualitative comparison

Five families, three questions

How cost grows with sequence length, how much of the model runs per token, and what that means for sizing. Curves show shape only.

Dense Transformer

Llama-3 70B, Mistral 7B

O(n²) attention
Total = active (100 %)

Memory proportional to full param count. Predictable but expensive at scale. KV cache grows linearly with batch × context.

Sparse MoE

Mixtral 8x7B, DeepSeek-V3

O(n²) attention
~3–30 % active

Must load all params onto device(s). Needs expert parallelism. FLOPs/token low but memory footprint is the total - not the active - parameter count.

SSM / Mamba

Mamba-2, Falcon Mamba

O(n) inference
Total = active (100 %)

No KV cache; fixed-size hidden state instead. Memory flat w.r.t. context length - transformative for very long sequences.

Hybrid (attention + linear / SSM)

Qwen3.5, Kimi K3, Nemotron 3, Jamba

Mixed O(n²) / O(n)
Dense or MoE backbone

KV cache only for the full-attention layers (1 in 4 to 1 in 10), plus a fixed-size state per sequence. Much less memory at long context, but the state is harder to prefix-cache and offload.

Multimodal VLM

Qwen3-VL, Llama 4 Maverick, InternVL3

O(n²) - images inflate n
Dense or MoE backbone

Image tokens occupy sequence budget. High-res inputs produce 500–2000 extra tokens per image - KV cache and prefill cost spike accordingly.

Gauge: solid = share of weights active per token; lighter tail = the range across models; striped = depends on the backbone.

Parameter counts and active-param ratios are illustrative; exact figures vary by configuration and version. Confirm against model cards and provider documentation before sizing production clusters.

Dig deeper

Architecture determines memory layout, serving topology, and cost - explore these pages to put the numbers on the table.