Applications
Agentic AI
AdvancedA chatbot answers a question. An agent pursues a goal - it plans, uses tools to act on the real world, observes what happened, and loops until it's done. This page builds from what 'agentic' actually means, through tools, memory, and orchestration, to why long-running agents are an infrastructure problem.
What makes a system “agentic”
The difference isn't the model - it's the autonomy. A single prompt is one input and one output. An is given a goal and the freedom to decide the steps: which tools to call, in what order, and when the job is finished. That loop is what turns a text generator into something that gets things done.
One prompt vs one goal
The same kind of request, handled two ways. Step through and watch what each side does - and what it still remembers at the end. Tool names are illustrative.
Single prompt (chatbot)
What it remembers
nothing yet
Agent
loop 1
What it remembers
nothing yet
The core: the reason–act–observe loop
Almost every agent runs the same loop, often called (reason + act). The model reasons about what to do next, acts by calling a tool, observes the result, and feeds that back in to reason again - repeating until the goal is reached. Step through a real run below.
The agent loop - reason, act, observe, repeat
A single prompt answers once. An agent circles a loop: it reasons about the goal, calls a tool outside the model, observes the result, and checks whether the goal is met - looping again or exiting. Example task: book a project kickoff.
Goal: book a kickoff for the new project. First I need everyone's availability.
Context window
4 blocks so far. Nothing is removed - the model re-reads all of it on the next step, which is the cost explored further down this page.
Tool use & function calling - how agents touch the world
On its own a model can only write text. (also called function calling) give it hands: a tool is a function the model can request - search the web, run code, query a database, send an email. The model outputs a structured call, your system runs it, and the result comes back as the next observation.
MCP - the USB-C port for AI tools
Before, every agent needed custom glue for every tool - an N×M integration mess. The , open-sourced by Anthropic and now adopted by OpenAI, Google, Microsoft, and AWS, makes it N+M: each agent speaks MCP, each tool exposes an MCP server, and any agent can discover and use any tool with no new code. An agent asked to “organize the kickoff” can hit a calendar server, an email server, a room-booking server, and a tracker - all in one reasoning session.
N×M glue code vs one protocol
Every line is an integration someone has to write and maintain. Add agents or tools and compare how fast each side grows.
Custom glue: N × M
12
integrations
With MCP: N + M
7
integrations
One MCP client per agent, one MCP server per tool: adding a tool is one new server that every agent can discover.
MCP SDK downloads / mo
97M+
Public MCP servers
10,000+
Integration cost
N+M
Figures as of early 2026. The Nov-2025 spec added async operations, server identity, and audit trails - the governance pieces enterprises were waiting for.
Context & memory - what the agent carries
An agent needs to remember what it has already done, learned, and been told. That memory comes in layers - fast working memory in the context window, and durable long-term memory it can look up when needed.
Where an agent keeps what it knows
Working memory sits next to the model; long-term memory lives outside it and is looked up when needed. Tap a layer. Sample entries are illustrative.
Long-term memory - durable, outside the model
Short-term (context window). The working memory of the current run - the system prompt, tool definitions, and everything reasoned and observed so far. Fast, but finite and re-read every step.
The catch: the only grows. Every step appends more reasoning and more tool output, and the model re-reads it all each time. That's the hidden cost driver - see it move below.
Why agents are expensive - context only grows
Every step appends the model's reasoning and the tool's result, and the model reads the whole context again before its next move. Without caching, that re-reading is recomputed from scratch each step, so total work grows much faster than the context itself.
Context read at each step
One column per step, fixed scale. All of it is computed again.
Running total of tokens computed
Sum of the computed part of every column so far.
Context at step 1
3.4K tok
Computed, cache off
3.4K tok
Computed, cache on
3.4K tok
Over 1 step, running with no cache computes 1.0× the tokens of running with a prefix cache. The context grows by a fixed amount per step, but the uncached total adds a whole column each step, so it bends upward. With a prefix cache, the unchanged system prompt, tool definitions, and earlier history stay in the KV cache - they are not recomputed, though the model still attends to them - so each step pays only for its new tokens.
Illustrative sizes: system prompt 600, tool definitions 1,800, and 250 reasoning + 700 tool-result tokens per step.
Multi-agent orchestration
One agent is sequential and can drift on a big task. The fix is to coordinate several - splitting work, running threads in parallel, and gating risky actions behind a human. These are the common patterns.
Four ways to organize agents
Pick a pattern and watch its messages flow. Several packets moving at once means work running in parallel. Example tasks are illustrative.
Orchestrator–worker
Example: a research question split into three threads
The goal goes to a lead agent.
Agents
4
Parallel now
1
Messages
1
A lead agent decomposes the goal and spawns sub-agents to work threads in parallel. Anthropic's research system beat a single agent by ~90% on internal evals this way.
Breadth-first work (research, search, review) finishes in fewer rounds because the threads overlap in time.
Where agents are working today
Agentic systems are already in production wherever a task is multi-step and needs to touch real tools.
Four jobs, one shape
Each production pattern is a multi-step loop. Steps marked touch a real tool; stacked steps fan out across several; the dashed line is where the agent loops back.
Coding agents
- read the repo
- plan a change
- edit files
- run tests
failing tests send it back to edit - it iterates on the errors
Cursor, Replit, and Claude Code read a repo, plan a change, edit files, run tests, and iterate on the errors - a loop, not a single completion.
Research agents
- decompose the question
- search many sources
- read + cross-check
- synthesize a cited report
cross-checking can send it back to search
Decompose a question, search many sources in parallel, read and cross-check, and synthesize a cited report.
Operations & RPA
- triage the ticket
- look up the customer
- fix across systems
- update the record
each system's result feeds the next step
Triage a ticket, look up the customer, take the fix across several systems, and update the record - the kickoff-booking demo, generalized.
Data & analytics
- question → queries
- run them
- check the numbers
- explain in plain language
checking the numbers can send it back to the queries
Translate a business question into queries, run them, check the numbers, and explain the result in plain language.
Agents are an infrastructure problem
A single chat is one model call. An agent is dozens - each re-reading a growing context, waiting on tools, and often retrieving from memory. That makes inference speed, reuse, and fast data access the difference between an agent that's magical and one that's too slow and too expensive to ship. And an agent's durable memory - its episodic, semantic, and vector stores - has to live somewhere persistent and instantly searchable: a data platform, not a transient context window.