skip to content

The layers of an AI agent product

/ 11 min read

Since agents became a thing, I decided that I wanted to build my own. I’m a believer that building your own will teach you more about the technology and make you a better user.

But. What is an “agent”?

“Agent” gets used for everything from a single API call to a full SaaS product. That makes it hard to talk about where a problem lives, or which piece of the stack should solve it. It helps to split an agent product into layers, each with its own job.

To keep it concrete, I’ll use one running example throughout: a support agent that can look up orders and issue refunds.

The structure

An AI agent product has distinct layers:

model SDK → agent framework → harness → execution runtime → product

Databases and files store state for several of these layers. Proof, meaning observability and evals, isn’t a layer: it measures how well all of them work.

LayerJobExamples
Model SDKOne request, one responseOpenAI SDK, Anthropic SDK, Google Gen AI SDK, Vercel AI SDK Core
Agent frameworkRun the tool-calling loopPydantic AI, Mastra, LangChain / LangGraph, OpenAI Agents SDK
HarnessDefine the job, context, and boundariesClaude Code, Codex CLI, GitHub Copilot CLI, Cursor, your own agent code
Execution runtimeRun the work durably, survive failuresTemporal, Inngest, DBOS, Celery + a queue, LangGraph checkpointers
ProductUsers, tenants, UI, billing, reconnectingChatGPT, Claude.ai, Cursor, Copilot coding agent, your SaaS
Proof (across all layers)Show what happened and whether it was goodOpenTelemetry, Langfuse, LangSmith, Logfire, Braintrust, Pydantic Evals

Model SDK

A model SDK makes a request and returns a response. That response might contain text or a request to call a tool, but the SDK does not complete a multi-step task on its own.

Here’s what that looks like with the Anthropic Python SDK:

import anthropic
client = anthropic.Anthropic()
tools = [{
"name": "get_order",
"description": "Look up an order by ID",
"input_schema": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
}]
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
tools=tools,
messages=[{"role": "user", "content": "Where is order 1234?"}],
)
print(response.stop_reason) # "tool_use"

The model didn’t look anything up. It returned a tool_use block that says “please call get_order with order_id="1234"”. Running that function, sending the result back, and asking again is your job. The SDK is done.

Agent framework

An agent framework manages that loop:

  1. Assemble context.
  2. Call the model.
  3. Validate the requested tool arguments.
  4. Execute the tools.
  5. Add their results to the conversation.
  6. Call the model again, until it finishes.

Stripped down, this is what’s under the hood of every framework:

messages = [{"role": "user", "content": "Where is order 1234?"}]
while True:
response = client.messages.create(
model="claude-sonnet-4-5", max_tokens=1024, tools=tools, messages=messages
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
break
results = []
for block in response.content:
if block.type == "tool_use":
output = TOOLS[block.name](**block.input) # validate first in real code
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": str(output),
})
messages.append({"role": "user", "content": results})

A framework gives you that loop plus the tedious parts: turning function signatures into JSON schemas, validating arguments, retrying when the model sends bad ones, streaming, and switching model providers. The same agent in Pydantic AI:

from pydantic_ai import Agent
agent = Agent("anthropic:claude-sonnet-4-5")
@agent.tool_plain
def get_order(order_id: str) -> dict:
"""Look up an order by ID."""
return orders.get(order_id)
result = agent.run_sync("Where is order 1234?")
print(result.output)

Other frameworks make different trade-offs:

  • Pydantic AI: Python, type-first. Tool arguments and outputs are validated with Pydantic models.
  • Mastra: TypeScript. Agents, tools with Zod schemas, and workflows in one package.
  • LangChain / LangGraph: LangGraph models the loop as an explicit graph of nodes and edges, which helps when the flow has branches.
  • OpenAI Agents SDK: A small loop with handoffs between agents and guardrails.

An agent run may make many model calls; it is not one continuous model request. Our “where is my order” question is already two calls: one that asks for the tool, one that writes the answer.

Harness

The framework is the loop. The harness is the job.

A framework answers how an agent runs: call the model, validate tool arguments, execute tools, append results, call again until it stops. It also covers the tedious parts around that loop: schemas, retries, streaming, provider switching. It doesn’t know what the agent is for.

A harness answers what this agent is allowed to be: instructions, which tools exist, what context is loaded, memory, permissions, budgets, and stopping rules. It determines what information reaches the model and what actions the model is allowed to request.

Think of a web app. FastAPI is the framework. Your routes, auth rules, and queries are the harness. FastAPI will route a request. It will not decide that refunds over $100 need a human.

QuestionFramework answersHarness answers
How do I call the model and run tools?The loop, schemas, validation, retriesn/a
Which tools exist?How to register themget_order and refund, nothing else
What does the model see?Where to put instructions and messagesThis customer’s summary, not the full history
What is it allowed to do?Gives you hooks for checksRefunds over $100 go to a human
When does it stop?Stops when the model stops calling toolsAfter 10 model calls, or when the budget runs out

The swap test makes it concrete. Swap the framework, port the same instructions, tools, and limits, and the agent should still do the same job. Swap the harness and you have a different agent, even on the same library. The same Pydantic AI loop can run a support agent that only reads one customer’s orders, or an analytics agent with read access to the whole warehouse. Same loop, different harness, different product.

A note on wording: products often call the whole thing a harness, loop included. Anthropic, for example, describes Claude Code as the agentic harness around Claude. Here, framework means the reusable loop, and harness means the job-specific decisions on top of it.

For our support agent, the harness is the code around agent.run():

from pydantic_ai import Agent, RunContext
from pydantic_ai.usage import UsageLimits
agent = Agent(
"anthropic:claude-sonnet-4-5",
deps_type=Customer,
instructions="You are a support agent. Only discuss the customer's own orders.",
)
@agent.instructions
def customer_context(ctx: RunContext[Customer]) -> str:
# Context selection: load a summary, not the whole history
return f"Customer: {ctx.deps.name}. Notes: {ctx.deps.support_summary}"
@agent.tool
def get_order(ctx: RunContext[Customer], order_id: str) -> dict:
# Scoping: the model can only see this customer's orders
return db.get_order(order_id, customer_id=ctx.deps.id)
@agent.tool
def refund(ctx: RunContext[Customer], order_id: str, amount: float) -> str:
# Permissions: large refunds need a human
if amount > 100:
return "Refund queued for human approval."
return payments.refund(order_id, amount)
result = agent.run_sync(
message,
deps=customer,
usage_limits=UsageLimits(request_limit=10), # Budget and stopping rule
)

Nothing in there is new framework machinery. It’s decisions: what the agent is for, what it can see, what it can do, and when it must stop.

Coding agents are the best-known harnesses. Claude Code, Codex CLI, and GitHub Copilot CLI all run roughly the same loop against a model. What makes them different is what’s under the hood of the harness:

  • Instructions: a long system prompt describing how to work in a codebase, plus project files such as CLAUDE.md or AGENTS.md.
  • Tools: read, edit, and search files; run shell commands; fetch URLs; spawn sub-agents.
  • Permissions: ask before running a command or editing a file, allow-lists for safe commands, sandboxes that block network or writes outside the project.
  • Context management: search the codebase instead of loading it, truncate large tool outputs, and summarise (compact) the conversation when it approaches the context limit.
  • Stopping rules: finish when the model stops calling tools, when the user interrupts, or when a limit is hit.

Swap the model under Claude Code and you still have most of Claude Code. Swap the harness and you have a different product. That’s also why Anthropic ships the Claude Agent SDK: it’s the Claude Code harness offered as a library.

Memory is part of that context strategy. It might be recent messages in a database, a summary, or a file such as MEMORY.md. The file is not inherently “memory”: the harness must decide when to write it, how to scope it, and what to load into future model requests. The model sees the context it is given, not everything stored on disk.

In the support agent above, support_summary is memory. Some job has to write it after each conversation, and customer_context decides to load it. Without both, it’s just a column in a table.

Execution runtime

The framework can run inside an HTTP request, but a durable service separates accepting a task from executing it. An API records a run; a worker executes it; checkpoints record completed progress so another worker can continue after a failure.

POST /runs → insert run (status=queued) → 202 Accepted { run_id }
↓
queue → worker → agent loop → checkpoint after each step
↓
run status=completed, events stored

The options range from simple to heavy:

  • A queue and workers (Celery, BullMQ, SQS + Lambda): easy to start. Checkpointing is up to you.
  • Durable execution engines (Temporal, Inngest, DBOS): each model call and tool call becomes a recorded step. If a worker dies, another one replays the recorded results and continues from the next step. Pydantic AI has built-in integrations for Temporal, DBOS, and Prefect.
  • Framework checkpointers: LangGraph saves graph state after each step to Postgres or SQLite, so a thread can resume or pause for human approval. Mastra workflows can suspend and resume.

What a durable execution engine actually is

An event log plus replay. Every engine has the same parts:

  • Storage: an append-only log per run (“step started”, “step completed with this result”), usually in Postgres.
  • A queue and workers: your code runs in your workers, built with the engine’s SDK. Temporal works like BullMQ with Redis: workers pull tasks from the Temporal server, which you self-host or rent as Temporal Cloud.
  • Replay: when a worker dies, another one re-runs the code from the top. Finished steps return their recorded result instead of running again, so execution continues from the first unfinished step. That’s the difference from a plain job queue, which retries the whole job.
  • Retries and timeouts: per step, with backoff, plus heartbeats so a dead worker’s task gets picked up.
  • Timers and signals: a run can wait days for a human approval without a process sleeping that long.

A basic version is a steps table keyed by (run_id, step_id) and a helper that checks it before running each step. That’s roughly how DBOS works, as a library on Postgres. Once you need timers, signals, and deploys while runs are in flight, use an engine.

A checkpoint is different from memory. Memory helps the agent answer the next question; a checkpoint says which steps of the current task have finished. It cannot freeze an in-flight model call or undo an email already sent. Retried actions therefore need idempotency or a way to reconcile their outcome.

For the refund tool, that means passing an idempotency key derived from the run and the step, so a retry after a crash doesn’t refund twice:

payments.refund(order_id, amount, idempotency_key=f"{run_id}:{step_id}")

Payment APIs such as Stripe support this directly. For APIs that don’t, check whether the action already happened before repeating it.

A database can store memory, checkpoints, and product data. That does not make “database” a separate agent layer. It is infrastructure serving different responsibilities.

Product and proof

The product adds authentication, tenancy, a UI, billing, and a way to reconnect to a run’s events. A browser connection should not own the work, and it should not be required for the run to complete or be charged.

The Copilot coding agent and Codex in the cloud are good examples: you assign a task, close the tab, and come back to a pull request. The run lives in the runtime, not in your browser. The UI just subscribes to its events, usually over SSE, and picks up where it left off on reconnect.

Observability shows what happened across model calls, tools, and workers. In practice that’s traces: one span per model call and tool call, with tokens, latency, and cost attached. OpenTelemetry has semantic conventions for generative AI, and tools such as Langfuse, LangSmith, and Logfire display them.

Evals measure whether the agent completed useful tasks at acceptable quality and cost. For the support agent, that’s a set of real conversations with expected outcomes: did it refund the right order, did it refuse to show another customer’s data, did it escalate the $500 refund? Run them on every prompt or model change with Braintrust, Pydantic Evals, or a plain test suite.

Tests and load measurements address operational reliability. Kill a worker mid-run. Deploy during a run. Send a thousand runs at once.

A run can finish reliably and still give a poor answer. It can also give an excellent answer once and fail to survive the next deploy. Both problems need measuring.