Applications

RAG & Vector Search

Intermediate

A trained model only knows what it saw before its cutoff. Retrieval-Augmented Generation fixes that: look up the right facts from your own data, then let the model answer with them in hand. Under the hood it is embeddings + vector search - turning meaning into numbers and search into geometry. This page builds from why RAG exists, to the real cosine numbers behind it, to billion-scale infrastructure.

Why RAG exists

A language model is a closed-book exam: it answers from memory frozen at its training cutoff. Ask about yesterday's news, an internal wiki, or a customer's contract and it has three options - say it doesn't know, or worse, confidently make something up (a hallucination). turns it into an open-book exam: before answering, the system retrieves the most relevant passages from a trusted source and pastes them into the prompt. The model then answers from the page in front of it, with citations.

Closed book vs open book

The same question, answered from memory alone and then with retrieval first. Tap a problem to see where in the picture it comes from.

Question

“What does clause 7 of our supplier contract say?”

Retrieve

Skipped. The private source is never consulted.

contract.pdfwikiticket

Model

Answers from memory frozen at its training cutoff.

trainingafter cutoff: ?

Answer

“I don't know” - or, worse, a confident, plausible guess (no source)

Knowledge cutoff. Models don't know anything after training. RAG injects fresh, current facts at query time.

The one-line mental model: RAG = look it up first, then answer.

Embeddings: turning meaning into numbers

To “look something up” by meaning, you first have to make meaning measurable. An model reads a piece of text (or an image) and outputs a list of numbers - a vector - that captures what it's about. The magic property: texts with similar meaning produce vectors that sit close together in that space, even when they share no words. So “car” and “automobile” land near each other; “car” and “banana” land far apart.

How long is the list? It depends on the model. Below are the dimension counts you'll meet most often.

Embedding modelDimensions (bar: 0 to 3072)Notes
OpenAI text-embedding-3-small1536default API embedder; Matryoshka-shortenable
OpenAI text-embedding-3-large3072~64.6% MTEB; highest quality OpenAI
Cohere embed v31024strong multilingual + reranker pairing
BGE-large / MPNet1024 / 768top open-weight encoders
all-MiniLM-L6-v2384tiny, fast, runs in the browser

To compare two vectors we use : the cosine of the angle between them, from −1 (opposite) through 0 (unrelated) to 1 (identical meaning). For example cos(king, queen) ≈ 0.86 (related) versus cos(king, banana) ≈ 0.12 (unrelated). The geometry is so faithful that vector arithmetic works: king − man + woman ≈ queen. Pick a query in the demo and watch the real numbers.

Search is geometry: similarity is an angle

Each word is a toy 6-D vector. The query points straight up and every other word sits at its true angle from it - cosine similarity is the cosine of that angle, so nearest-neighbor search returns the k smallest angles.

Query
top-k
0.90.750.50.25030°query: kingscale = cosine · left or right has no meaningqueenmanwomandoctornursetruckcarapplebanana
royaltypeoplefruitvehicletop-3 cone (±61°)

Ranked by cosine to “king”

cos(A, B) = A·B / (‖A‖ ‖B‖) = cos θ

cos(king, queen) = 0.863 → θ = 30.3°

Dividing by the lengths throws magnitude away - only direction counts. 1 = same direction, 0 = unrelated (90°), −1 = opposite.

For king, the nearest word is queen (cos 0.86) and the farthest is banana (cos 0.12) - no keywords matched, only angles. That is the whole idea behind vector search: similar meaning points the same way, so retrieval becomes a smallest-angle lookup.

Toy vectors, hand-tuned for teaching; real embeddings have hundreds to thousands of dimensions. On these same vectors, king − man + woman lands nearest queen (cos 0.98).

Vector search fundamentals

Once everything is a vector, retrieval is just nearest-neighbor search: embed the query, then find the stored vectors closest to it. The first question is how you measure “close.”

Three ways to measure “close”

Two vectors, three metrics. Turn the angle and stretch b, then normalize and watch the three agree. Drawn in 2-D for illustration; real embeddings have hundreds of dimensions.

length 1θab

cosine reads only the angle θ

Cosine similarity. Angle between vectors, ignoring length. Range −1..1. The default for text embeddings because magnitude carries little meaning.

Stretch b and only the dot product and L2 move - cosine stays 0.77 because it ignores length.

The second question is how you find those neighbors without checking all billion vectors every time. That is the difference between exact and approximate search.

Exact vs approximate: four index strategies

One query (◆) against a toy set of 48 vectors. Count how many each strategy has to compare to find the three nearest. Counts are from this toy set, for illustration.

query

compared

48 of 48

result

top 3 found: 3 of 3 (exact)

Brute-force kNN (exact). Compare the query to every vector. 100% recall, but O(N·d) per query - only viable below roughly 1M vectors.

true 3 nearest

indexes give up a tiny bit of accuracy to go orders of magnitude faster - typically 95–99% recall at sub-10ms. The demo below lets you feel that tradeoff at different corpus sizes.

The recall ↔ latency ↔ cost triangle

Exact search is always 100% accurate but scans every vector. ANN walks a graph and touches only a few, trading a sliver of recall for orders-of-magnitude speed. Grow the corpus, turn the index knob and compress the vectors - the shape bends. Illustrative model.

Stored as

Pick two corners

outer edge = best

RecallANN 97.5%·exact 100%SpeedANN 4.8 msexact 60 msLow memoryANN 6.1 GBexact 6.1 GB
ANN index (float32)exact brute force (float32)

Why ANN is fast

entryquery
4/64 compared

Hop 1: jump to the best node so far and check its neighbors · 4 checked

1/12

Toy 2-D data, 64 points. Exact work grows with N; a graph walk grows roughly with log N - that gap is why ANN wins at scale.

ANN recall

97.5%

exact = 100%

Query latency

4.8 ms

exact 60 ms · ANN ~13× faster

Index memory (1536-d)

6.1 GB

float32 6.1 GB → float32

Under ~1M vectors, exact is still fast enough. 1536 × 4 B = 6 KB per float32 vector.

1B × 1536 dims × 4 B = ~6.1 TB float32 → int8 ~1.5 TB → PQ ~tens–hundreds of GB

Pick any two corners of the triangle and the third gives way. Billion-scale production indexes lean on ANN + quantization, and GPU-accelerated builds (NVIDIA cuVS / CAGRA) to make the index in hours instead of days.

The RAG pipeline, end to end

RAG has two halves. Offline (ingest): documents are split into (~200–500 tokens with 10–20% overlap so ideas aren't cut mid-sentence), each chunk is embedded, and the vectors are stored in the index. Online (query): the question is embedded, the top-k chunks (usually 3–20) are retrieved, an optional - a cross-encoder that reads query and chunk together - reorders them for precision, and the survivors are stuffed into the prompt for the LLM to answer with citations.

One query, end to end

Step through the online half of RAG: embed, retrieve, rerank, build the prompt, generate. Chunks and scores are illustrative.

1/6A user question arrives as plain text.

Query in: A user question arrives as plain text.

1. Query in

“How does the KV cache lower inference cost?”

≈ 9 tokens

No keyword index is consulted. The question itself is about to become a point in the same vector space as every chunk.

Chunks in the index (not retrieved yet)

c2 · serving-notes.md #7

Decode is memory-bandwidth bound; reusing cached KV avoids recomputing past tokens at every step.

140 tokens

c1 · kv-cache-guide.md #3

The KV cache stores the keys and values of every past token so they are not recomputed each step.

120 tokens

c3 · gpu-overview.md #2

HBM bandwidth and VRAM capacity set the ceiling on batch size and context length for a given GPU.

110 tokens

Prompt tokens (prefill)49 / 419

system 40 + query 9 + chunks not added yet

Vector search ranked the serving note first, but the reranker reads each chunk against the question: the KV-cache guide moves up and the loosely related GPU chunk drops from 0.79 to 0.55. Better retrieval means fewer, sharper chunks, which directly lowers prefill cost and KV-cache size downstream.

Why an ML engineer cares about the token meter: every retrieved chunk lengthens the prompt, which lengthens prefill and grows the KV cache. Better retrieval - fewer, sharper chunks - is not just more accurate, it is cheaper to serve.

The technology landscape

A RAG stack is assembled from four kinds of building block. Here are the names you'll hear in vendor calls - what each one is, and why it matters.

Where each building block sits in a request

A question flows left to right; GPU acceleration sits underneath the vector search. Tap a block for the names you'll hear in vendor calls.

LLM

the survivors go into the prompt

Vector databases

  • Pinecone - Fully-managed serverless vector DB; zero-ops scaling.
  • Weaviate - Open-source DB with built-in hybrid search & modules.
  • Qdrant - Rust-based, fast HNSW with payload filtering.
  • Milvus - Cloud-native, billion-scale; GPU indexing via cuVS.
  • pgvector - Postgres extension - vectors next to your relational data.
  • FAISS - Meta's library; the de-facto ANN engine many DBs embed.

Scale & cost

Every production decision lives inside one triangle: recall ↔ latency ↔ cost. You can have any two; the third bends. The dominant cost driver is memory, because the highest-recall index () wants the whole graph resident in RAM.

When to use what - by corpus size

Drag the corpus size along the log axis; the lanes it crosses are the tools that fit. Band edges are approximate.

32M vectors
Exact brute force or pgvector
HNSW or IVF + quantization
GPU ANN builds (cuVS / CAGRA)

1M–1B+ vectors

HNSW or IVF + quantization

ANN p99 in single-digit to low-tens of ms.

Exact search stops scaling first; past ~1M vectors an ANN index takes over, and at billion scale GPU-accelerated builds keep index construction to hours.

Hyperscale math: what a billion (or trillion) vectors costs

The reason “just put it all in RAM” stops working is arithmetic: an HNSW graph adds links on top of the raw vectors, while quantization shrinks them. Size an index below and watch the total cross the memory tiers - that crossover is why a trillion-vector index can live on NVMe but not in DRAM.

What a billion (or trillion) vectors actually costs

Raw footprint is num_vectors × dimensions × bytes/component. The index then piles structure on top (HNSW) or shrinks the payload (PQ). Drag the controls and watch the total cross the host-RAM wall. Units are decimal (1 GB = 10⁹ B).

Component precision
Index type

HNSW - recall very high. Graph index - highest recall & speed, but RAM-hungry (graph + raw).

Total footprint - HNSW

6.89 TB

GPU HBMHost RAMLocal NVMeVAST pool
HNSW total
6.89 TB
PQ
doesn't fit
doesn't fit
fits
fits
1 GB
1 TB
1 PB
1 EB

Host RAM (DRAM) - where HNSW/IVF indexes must sit in classic ANN libraries: hundreds of GB - low TB, ~100 ns access. Tap a tier for its role.

What the total is made of

Raw vectors × 1.08 slack 6.64 TB (96%)HNSW graph links (M=32) 256.0 GB (4%)

1.0B vectors × 1536 dims in FP32 is 6.14 TB of raw payload. HNSW adds a 256.0 GB graph on top, and the whole index must live resident in RAM to stay fast. That is past the host-RAM band, so keeping it in memory means sharding across servers.

Show the math›

The recall ↔ latency ↔ cost triangle. HNSW buys the highest recall and lowest latency by keeping a big graph resident in RAM. IVF and PQ cut that cost - IVF by probing only a few lists, PQ by compressing the vectors themselves - at the price of recall. VAST's hierarchical, NVMe-resident index sidesteps the “whole index must fit in RAM” constraint these in-memory libraries assume.

Estimates use standard rules of thumb: HNSW graph ≈ N × M × 2 × 4 B (M=32) plus ~8% allocator slack on the raw vectors; IVF nlist ≈ √N (capped); PQ m=16, k=256 (8-bit codes). Real systems vary with build parameters - treat these as order-of-magnitude. Tier bands are order-of-magnitude media classes, not VAST-specific capacities.

Use case: when and where to position RAG

Theory lands when you can see it on a real desk. Here's a composite enterprise scenario - a large insurer's claims-and-service copilot - that shows the moment RAG earns its place: the same model answers a customer question with and without retrieval over the company's own corpus, data lake, and data warehouse. Toggle it and watch a plausible-but-wrong answer become a grounded, citable one.

In the wild - a P&C insurer's claims & service copilot

Contact-center agents and adjusters were drowning in policy language and giving inconsistent, sometimes wrong answers. RAG over their own data turns a generic chatbot into a grounded, auditable expert: every claim in the answer traces back to a source. Tap a claim to see where it came from.

Illustrative
Question
Retrieval

“A policyholder's basement flooded after a pipe burst. Is it covered, and what's their deductible?”

Retrieved from the data foundation

Enterprise corpus

HO-3 wording ¶12 · endorsement HO-04-95

Data lake

FNOL transcript · first adjuster note

Data warehouse

deductible $1,500 · coverage limits · form on file

Live status

claim #CLM-48217 → status: inspection

Grounded answer · every claim linked

Yes - sudden, accidental discharge from a burst pipe1 is covered under Section I, Coverage A of the HO-3 form (¶12, “Water Damage”).2 This policyholder also carries the Water Backup & Sump Overflow endorsement (HO-04-95), adding $10,000 of coverage.3 Their all-perils deductible is $1,500 - not the $2,500 wind/hail deductible.4 External surface flooding remains excluded.5 The claim is open and currently in “inspection.”6

Underline color = source. Plain text is the copilot's recommendation, not a retrieved fact.

Outcome after rollout

before after

Average handle time−38%

11.5 min → 7.1 min

First-contact resolution+14 pts

66% → 80%

New-adjuster ramp−50%

6 mo → 3 mo

Incorrect coverage statements (QA)−72%

baseline → -72% (index)

The point: the model didn't get smarter - the data got closer. Answers became specific, correct, and auditable because every claim is grounded in the company's own corpus, lake, and warehouse. That is why, at enterprise scale, RAG is a data-platform decision, not just a prompt.

Composite, illustrative scenario; figures are representative of published enterprise RAG deployments, not a specific customer.

The separate vector DB problem

Notice what the stack above quietly assumes: the embeddings live in a dedicated - Pinecone, Milvus, Weaviate, or pgvector - that is separate from where the source records actually live (your data lake, warehouse, or object store). That split is the default today, and it quietly creates work.

One record, two copies

Follow a single policy document as it changes. The record lives in the system of record; its vector lives in a separate vector database.

1/4Two systems: the records in one, their vectors in another.

System of record

data lake · warehouse · object store

policy-42 v1current

permission model Acluster 1

copy + ETL copy embed loadidle

Vector database

Pinecone · Milvus · Weaviate · pgvector

vec(policy-42 v1)in sync
query returns v1

permission model Bcluster 2

The obvious question is whether vectors could just live next to the data they describe - one system, no copy, no drift. That is exactly the step the VAST VectorStore page picks up.

Retrieval is upstream of everything you serve

The chunks you retrieve become the prompt you serve, so RAG quality flows straight into prefill cost and KV-cache size. And at scale, vector search is a data-platform problem - which is exactly where VAST VectorStore runs search native to the data it lives on.