Applications
RAG & Vector Search
IntermediateA trained model only knows what it saw before its cutoff. Retrieval-Augmented Generation fixes that: look up the right facts from your own data, then let the model answer with them in hand. Under the hood it is embeddings + vector search - turning meaning into numbers and search into geometry. This page builds from why RAG exists, to the real cosine numbers behind it, to billion-scale infrastructure.
Why RAG exists
A language model is a closed-book exam: it answers from memory frozen at its training cutoff. Ask about yesterday's news, an internal wiki, or a customer's contract and it has three options - say it doesn't know, or worse, confidently make something up (a hallucination). turns it into an open-book exam: before answering, the system retrieves the most relevant passages from a trusted source and pastes them into the prompt. The model then answers from the page in front of it, with citations.
Closed book vs open book
The same question, answered from memory alone and then with retrieval first. Tap a problem to see where in the picture it comes from.
Question
“What does clause 7 of our supplier contract say?”
Retrieve
Skipped. The private source is never consulted.
Model
Answers from memory frozen at its training cutoff.
Answer
“I don't know” - or, worse, a confident, plausible guess (no source)
Knowledge cutoff. Models don't know anything after training. RAG injects fresh, current facts at query time.
The one-line mental model: RAG = look it up first, then answer.
Embeddings: turning meaning into numbers
To “look something up” by meaning, you first have to make meaning measurable. An model reads a piece of text (or an image) and outputs a list of numbers - a vector - that captures what it's about. The magic property: texts with similar meaning produce vectors that sit close together in that space, even when they share no words. So “car” and “automobile” land near each other; “car” and “banana” land far apart.
How long is the list? It depends on the model. Below are the dimension counts you'll meet most often.
| Embedding model | Dimensions (bar: 0 to 3072) | Notes |
|---|---|---|
| OpenAI text-embedding-3-small | 1536 | default API embedder; Matryoshka-shortenable |
| OpenAI text-embedding-3-large | 3072 | ~64.6% MTEB; highest quality OpenAI |
| Cohere embed v3 | 1024 | strong multilingual + reranker pairing |
| BGE-large / MPNet | 1024 / 768 | top open-weight encoders |
| all-MiniLM-L6-v2 | 384 | tiny, fast, runs in the browser |
To compare two vectors we use : the cosine of the angle between them, from −1 (opposite) through 0 (unrelated) to 1 (identical meaning). For example cos(king, queen) ≈ 0.86 (related) versus cos(king, banana) ≈ 0.12 (unrelated). The geometry is so faithful that vector arithmetic works: king − man + woman ≈ queen. Pick a query in the demo and watch the real numbers.
Search is geometry: similarity is an angle
Each word is a toy 6-D vector. The query points straight up and every other word sits at its true angle from it - cosine similarity is the cosine of that angle, so nearest-neighbor search returns the k smallest angles.
Ranked by cosine to “king”
cos(A, B) = A·B / (‖A‖ ‖B‖) = cos θ
cos(king, queen) = 0.863 → θ = 30.3°
Dividing by the lengths throws magnitude away - only direction counts. 1 = same direction, 0 = unrelated (90°), −1 = opposite.
For king, the nearest word is queen (cos 0.86) and the farthest is banana (cos 0.12) - no keywords matched, only angles. That is the whole idea behind vector search: similar meaning points the same way, so retrieval becomes a smallest-angle lookup.
Toy vectors, hand-tuned for teaching; real embeddings have hundreds to thousands of dimensions. On these same vectors, king − man + woman lands nearest queen (cos 0.98).
Vector search fundamentals
Once everything is a vector, retrieval is just nearest-neighbor search: embed the query, then find the stored vectors closest to it. The first question is how you measure “close.”
Three ways to measure “close”
Two vectors, three metrics. Turn the angle and stretch b, then normalize and watch the three agree. Drawn in 2-D for illustration; real embeddings have hundreds of dimensions.
cosine reads only the angle θ
Cosine similarity. Angle between vectors, ignoring length. Range −1..1. The default for text embeddings because magnitude carries little meaning.
Stretch b and only the dot product and L2 move - cosine stays 0.77 because it ignores length.
The second question is how you find those neighbors without checking all billion vectors every time. That is the difference between exact and approximate search.
Exact vs approximate: four index strategies
One query (◆) against a toy set of 48 vectors. Count how many each strategy has to compare to find the three nearest. Counts are from this toy set, for illustration.
compared
48 of 48
result
top 3 found: 3 of 3 (exact)
Brute-force kNN (exact). Compare the query to every vector. 100% recall, but O(N·d) per query - only viable below roughly 1M vectors.
true 3 nearest
indexes give up a tiny bit of accuracy to go orders of magnitude faster - typically 95–99% recall at sub-10ms. The demo below lets you feel that tradeoff at different corpus sizes.
The recall ↔ latency ↔ cost triangle
Exact search is always 100% accurate but scans every vector. ANN walks a graph and touches only a few, trading a sliver of recall for orders-of-magnitude speed. Grow the corpus, turn the index knob and compress the vectors - the shape bends. Illustrative model.
Pick two corners
outer edge = best
Why ANN is fast
Hop 1: jump to the best node so far and check its neighbors · 4 checked
Toy 2-D data, 64 points. Exact work grows with N; a graph walk grows roughly with log N - that gap is why ANN wins at scale.
ANN recall
97.5%
exact = 100%
Query latency
4.8 ms
exact 60 ms · ANN ~13× faster
Index memory (1536-d)
6.1 GB
float32 6.1 GB → float32
Under ~1M vectors, exact is still fast enough. 1536 × 4 B = 6 KB per float32 vector.
Pick any two corners of the triangle and the third gives way. Billion-scale production indexes lean on ANN + quantization, and GPU-accelerated builds (NVIDIA cuVS / CAGRA) to make the index in hours instead of days.
The RAG pipeline, end to end
RAG has two halves. Offline (ingest): documents are split into (~200–500 tokens with 10–20% overlap so ideas aren't cut mid-sentence), each chunk is embedded, and the vectors are stored in the index. Online (query): the question is embedded, the top-k chunks (usually 3–20) are retrieved, an optional - a cross-encoder that reads query and chunk together - reorders them for precision, and the survivors are stuffed into the prompt for the LLM to answer with citations.
One query, end to end
Step through the online half of RAG: embed, retrieve, rerank, build the prompt, generate. Chunks and scores are illustrative.
Query in: A user question arrives as plain text.
1. Query in
≈ 9 tokens
No keyword index is consulted. The question itself is about to become a point in the same vector space as every chunk.
Chunks in the index (not retrieved yet)
Decode is memory-bandwidth bound; reusing cached KV avoids recomputing past tokens at every step.
140 tokens
The KV cache stores the keys and values of every past token so they are not recomputed each step.
120 tokens
HBM bandwidth and VRAM capacity set the ceiling on batch size and context length for a given GPU.
110 tokens
system 40 + query 9 + chunks not added yet
Vector search ranked the serving note first, but the reranker reads each chunk against the question: the KV-cache guide moves up and the loosely related GPU chunk drops from 0.79 to 0.55. Better retrieval means fewer, sharper chunks, which directly lowers prefill cost and KV-cache size downstream.
Why an ML engineer cares about the token meter: every retrieved chunk lengthens the prompt, which lengthens prefill and grows the KV cache. Better retrieval - fewer, sharper chunks - is not just more accurate, it is cheaper to serve.
The technology landscape
A RAG stack is assembled from four kinds of building block. Here are the names you'll hear in vendor calls - what each one is, and why it matters.
Where each building block sits in a request
A question flows left to right; GPU acceleration sits underneath the vector search. Tap a block for the names you'll hear in vendor calls.
LLM
the survivors go into the prompt
Vector databases
- Pinecone - Fully-managed serverless vector DB; zero-ops scaling.
- Weaviate - Open-source DB with built-in hybrid search & modules.
- Qdrant - Rust-based, fast HNSW with payload filtering.
- Milvus - Cloud-native, billion-scale; GPU indexing via cuVS.
- pgvector - Postgres extension - vectors next to your relational data.
- FAISS - Meta's library; the de-facto ANN engine many DBs embed.
Scale & cost
Every production decision lives inside one triangle: recall ↔ latency ↔ cost. You can have any two; the third bends. The dominant cost driver is memory, because the highest-recall index () wants the whole graph resident in RAM.
When to use what - by corpus size
Drag the corpus size along the log axis; the lanes it crosses are the tools that fit. Band edges are approximate.
1M–1B+ vectors
HNSW or IVF + quantization
ANN p99 in single-digit to low-tens of ms.
Exact search stops scaling first; past ~1M vectors an ANN index takes over, and at billion scale GPU-accelerated builds keep index construction to hours.
Hyperscale math: what a billion (or trillion) vectors costs
The reason “just put it all in RAM” stops working is arithmetic: an HNSW graph adds links on top of the raw vectors, while quantization shrinks them. Size an index below and watch the total cross the memory tiers - that crossover is why a trillion-vector index can live on NVMe but not in DRAM.
What a billion (or trillion) vectors actually costs
Raw footprint is num_vectors × dimensions × bytes/component. The index then piles structure on top (HNSW) or shrinks the payload (PQ). Drag the controls and watch the total cross the host-RAM wall. Units are decimal (1 GB = 10⁹ B).
HNSW - recall very high. Graph index - highest recall & speed, but RAM-hungry (graph + raw).
Total footprint - HNSW
6.89 TB
Host RAM (DRAM) - where HNSW/IVF indexes must sit in classic ANN libraries: hundreds of GB - low TB, ~100 ns access. Tap a tier for its role.
What the total is made of
1.0B vectors × 1536 dims in FP32 is 6.14 TB of raw payload. HNSW adds a 256.0 GB graph on top, and the whole index must live resident in RAM to stay fast. That is past the host-RAM band, so keeping it in memory means sharding across servers.
Show the math›
The recall ↔ latency ↔ cost triangle. HNSW buys the highest recall and lowest latency by keeping a big graph resident in RAM. IVF and PQ cut that cost - IVF by probing only a few lists, PQ by compressing the vectors themselves - at the price of recall. VAST's hierarchical, NVMe-resident index sidesteps the “whole index must fit in RAM” constraint these in-memory libraries assume.
Estimates use standard rules of thumb: HNSW graph ≈ N × M × 2 × 4 B (M=32) plus ~8% allocator slack on the raw vectors; IVF nlist ≈ √N (capped); PQ m=16, k=256 (8-bit codes). Real systems vary with build parameters - treat these as order-of-magnitude. Tier bands are order-of-magnitude media classes, not VAST-specific capacities.
Use case: when and where to position RAG
Theory lands when you can see it on a real desk. Here's a composite enterprise scenario - a large insurer's claims-and-service copilot - that shows the moment RAG earns its place: the same model answers a customer question with and without retrieval over the company's own corpus, data lake, and data warehouse. Toggle it and watch a plausible-but-wrong answer become a grounded, citable one.
In the wild - a P&C insurer's claims & service copilot
Contact-center agents and adjusters were drowning in policy language and giving inconsistent, sometimes wrong answers. RAG over their own data turns a generic chatbot into a grounded, auditable expert: every claim in the answer traces back to a source. Tap a claim to see where it came from.
“A policyholder's basement flooded after a pipe burst. Is it covered, and what's their deductible?”
Retrieved from the data foundation
HO-3 wording ¶12 · endorsement HO-04-95
FNOL transcript · first adjuster note
deductible $1,500 · coverage limits · form on file
claim #CLM-48217 → status: inspection
Yes - sudden, accidental discharge from a burst pipe1 is covered under Section I, Coverage A of the HO-3 form (¶12, “Water Damage”).2 This policyholder also carries the Water Backup & Sump Overflow endorsement (HO-04-95), adding $10,000 of coverage.3 Their all-perils deductible is $1,500 - not the $2,500 wind/hail deductible.4 External surface flooding remains excluded.5 The claim is open and currently in “inspection.”6
Underline color = source. Plain text is the copilot's recommendation, not a retrieved fact.
Outcome after rollout
before after
11.5 min → 7.1 min
66% → 80%
6 mo → 3 mo
baseline → -72% (index)
The point: the model didn't get smarter - the data got closer. Answers became specific, correct, and auditable because every claim is grounded in the company's own corpus, lake, and warehouse. That is why, at enterprise scale, RAG is a data-platform decision, not just a prompt.
Composite, illustrative scenario; figures are representative of published enterprise RAG deployments, not a specific customer.
The separate vector DB problem
Notice what the stack above quietly assumes: the embeddings live in a dedicated - Pinecone, Milvus, Weaviate, or pgvector - that is separate from where the source records actually live (your data lake, warehouse, or object store). That split is the default today, and it quietly creates work.
One record, two copies
Follow a single policy document as it changes. The record lives in the system of record; its vector lives in a separate vector database.
System of record
data lake · warehouse · object store
permission model Acluster 1
Vector database
Pinecone · Milvus · Weaviate · pgvector
permission model Bcluster 2
The obvious question is whether vectors could just live next to the data they describe - one system, no copy, no drift. That is exactly the step the VAST VectorStore page picks up.
Retrieval is upstream of everything you serve
The chunks you retrieve become the prompt you serve, so RAG quality flows straight into prefill cost and KV-cache size. And at scale, vector search is a data-platform problem - which is exactly where VAST VectorStore runs search native to the data it lives on.