Applications

Agentic AI

Advanced

A chatbot answers a question. An agent pursues a goal - it plans, uses tools to act on the real world, observes what happened, and loops until it's done. This page builds from what 'agentic' actually means, through tools, memory, and orchestration, to why long-running agents are an infrastructure problem.

What makes a system “agentic”

The difference isn't the model - it's the autonomy. A single prompt is one input and one output. An is given a goal and the freedom to decide the steps: which tools to call, in what order, and when the job is finished. That loop is what turns a text generator into something that gets things done.

One prompt vs one goal

The same kind of request, handled two ways. Step through and watch what each side does - and what it still remembers at the end. Tool names are illustrative.

1/4Both get a request. The chatbot gets a question; the agent gets a goal.

Single prompt (chatbot)

QuestionModelAnswer (text)

What it remembers

nothing yet

Agent

Goal
ReasonActObserve

loop 1

Task done

What it remembers

nothing yet

In / outChatbot: One question in, one answer out.Agent: A goal in, a completed task out.
ActionsChatbot: No actions - it can only produce text.Agent: Calls tools to read and change the world.
KnowledgeChatbot: Knows only what's in the prompt and its training.Agent: Gathers new information as it goes.
StateChatbot: Stateless: forgets the moment the reply is done.Agent: Keeps state across many steps until the goal is met.

The core: the reason–act–observe loop

Almost every agent runs the same loop, often called (reason + act). The model reasons about what to do next, acts by calling a tool, observes the result, and feeds that back in to reason again - repeating until the goal is reached. Step through a real run below.

The agent loop - reason, act, observe, repeat

A single prompt answers once. An agent circles a loop: it reasons about the goal, calls a tool outside the model, observes the result, and checks whether the goal is met - looping again or exiting. Example task: book a project kickoff.

1/17
no - loop againyes: exitGoal met - exitModelloop 1ReasonActObserveGoalmet?EXTERNAL TOOLScalendarexternal toolroomsexternal tooltrackerexternal tool
ReasonLoop 1

Goal: book a kickoff for the new project. First I need everyone's availability.

Context window

systemtool defsgoalreason 1

4 blocks so far. Nothing is removed - the model re-reads all of it on the next step, which is the cost explored further down this page.

Tool use & function calling - how agents touch the world

On its own a model can only write text. (also called function calling) give it hands: a tool is a function the model can request - search the web, run code, query a database, send an email. The model outputs a structured call, your system runs it, and the result comes back as the next observation.

MCP - the USB-C port for AI tools

Before, every agent needed custom glue for every tool - an N×M integration mess. The , open-sourced by Anthropic and now adopted by OpenAI, Google, Microsoft, and AWS, makes it N+M: each agent speaks MCP, each tool exposes an MCP server, and any agent can discover and use any tool with no new code. An agent asked to “organize the kickoff” can hit a calendar server, an email server, a room-booking server, and a tracker - all in one reasoning session.

N×M glue code vs one protocol

Every line is an integration someone has to write and maintain. Add agents or tools and compare how fast each side grows.

Agents3
Tools4
MCPAgent 1Agent 2Agent 3calendaremailroomstracker

Custom glue: N × M

12

integrations

With MCP: N + M

7

integrations

One MCP client per agent, one MCP server per tool: adding a tool is one new server that every agent can discover.

MCP SDK downloads / mo

97M+

Public MCP servers

10,000+

Integration cost

N+M

Figures as of early 2026. The Nov-2025 spec added async operations, server identity, and audit trails - the governance pieces enterprises were waiting for.

Context & memory - what the agent carries

An agent needs to remember what it has already done, learned, and been told. That memory comes in layers - fast working memory in the context window, and durable long-term memory it can look up when needed.

Where an agent keeps what it knows

Working memory sits next to the model; long-term memory lives outside it and is looked up when needed. Tap a layer. Sample entries are illustrative.

Long-term memory - durable, outside the model

Short-term (context window). The working memory of the current run - the system prompt, tool definitions, and everything reasoned and observed so far. Fast, but finite and re-read every step.

The catch: the only grows. Every step appends more reasoning and more tool output, and the model re-reads it all each time. That's the hidden cost driver - see it move below.

Why agents are expensive - context only grows

Every step appends the model's reasoning and the tool's result, and the model reads the whole context again before its next move. Without caching, that re-reading is recomputed from scratch each step, so total work grows much faster than the context itself.

Prefix cache
1/30

Context read at each step

One column per step, fixed scale. All of it is computed again.

010K20K30K1102030Step 1: context 3,350 tokens; computed 3,3503.4K
SystemTool defsHistoryNew this step

Running total of tokens computed

Sum of the computed part of every column so far.

0250K500K1102030
Cache offCache on

Context at step 1

3.4K tok

Computed, cache off

3.4K tok

Computed, cache on

3.4K tok

Over 1 step, running with no cache computes 1.0× the tokens of running with a prefix cache. The context grows by a fixed amount per step, but the uncached total adds a whole column each step, so it bends upward. With a prefix cache, the unchanged system prompt, tool definitions, and earlier history stay in the KV cache - they are not recomputed, though the model still attends to them - so each step pays only for its new tokens.

Illustrative sizes: system prompt 600, tool definitions 1,800, and 250 reasoning + 700 tool-result tokens per step.

Multi-agent orchestration

One agent is sequential and can drift on a big task. The fix is to coordinate several - splitting work, running threads in parallel, and gating risky actions behind a human. These are the common patterns.

Four ways to organize agents

Pick a pattern and watch its messages flow. Several packets moving at once means work running in parallel. Example tasks are illustrative.

1/6
GoalLead agentplans and mergesAnswerWorker 1thread 1Worker 2thread 2Worker 3thread 3

Orchestrator–worker

Example: a research question split into three threads

The goal goes to a lead agent.

Agents

4

Parallel now

1

Messages

1

A lead agent decomposes the goal and spawns sub-agents to work threads in parallel. Anthropic's research system beat a single agent by ~90% on internal evals this way.

request / hand-offresult coming back

Breadth-first work (research, search, review) finishes in fewer rounds because the threads overlap in time.

Where agents are working today

Agentic systems are already in production wherever a task is multi-step and needs to touch real tools.

Four jobs, one shape

Each production pattern is a multi-step loop. Steps marked touch a real tool; stacked steps fan out across several; the dashed line is where the agent loops back.

Coding agents
  1. read the repo
  2. plan a change
  3. edit files
  4. run tests

failing tests send it back to edit - it iterates on the errors

Cursor, Replit, and Claude Code read a repo, plan a change, edit files, run tests, and iterate on the errors - a loop, not a single completion.

Research agents
  1. decompose the question
  2. search many sources
  3. read + cross-check
  4. synthesize a cited report

cross-checking can send it back to search

Decompose a question, search many sources in parallel, read and cross-check, and synthesize a cited report.

Operations & RPA
  1. triage the ticket
  2. look up the customer
  3. fix across systems
  4. update the record

each system's result feeds the next step

Triage a ticket, look up the customer, take the fix across several systems, and update the record - the kickoff-booking demo, generalized.

Data & analytics
  1. question → queries
  2. run them
  3. check the numbers
  4. explain in plain language

checking the numbers can send it back to the queries

Translate a business question into queries, run them, check the numbers, and explain the result in plain language.

Agents are an infrastructure problem

A single chat is one model call. An agent is dozens - each re-reading a growing context, waiting on tools, and often retrieving from memory. That makes inference speed, reuse, and fast data access the difference between an agent that's magical and one that's too slow and too expensive to ship. And an agent's durable memory - its episodic, semantic, and vector stores - has to live somewhere persistent and instantly searchable: a data platform, not a transient context window.