Building Type-Safe LLM Agents with PydanticAI
PydanticAI is an agent framework from the team behind Pydantic, designed to bring the same developer experience that FastAPI brought to web APIs into the world of generative AI. This tutorial walks through its core concepts: type-safe agents, structured outputs, dependency injection, and tools, ending with a complete customer-support agent you can adapt to your own projects.
What PydanticAI Is and Why It Exists
The Pydantic team built PydanticAI because much of the Python AI ecosystem either reinvented validation poorly or wrapped LLM calls in loosely typed dictionaries. PydanticAI takes a different stance: lean on the type system, validate everything, and keep the API small enough to reason about.
A few design principles shape the framework:
- Type-safe by design. Agents declare their dependency type and output type as generic parameters. Static checkers like mypy and Pyright understand them, and your editor autocompletes accordingly.
- Model-agnostic. OpenAI, Anthropic, Google (Gemini), Groq, Mistral, Ollama, and others are supported behind a single
Agentinterface. Switching models is usually a one-line change. - FastAPI-like ergonomics. Dependency injection, decorators for tools and validators, and a clean separation between defining an agent and running it will feel familiar if you have used FastAPI.
- Built on Pydantic. Structured outputs are plain Pydantic models, so you get validation, JSON schema generation, and serialization for free.
- Production observability. First-class integration with Logfire gives you tracing and debugging without bolting on a separate framework.
PydanticAI is intentionally not a do-everything orchestration platform. It focuses on the agent loop, leaving you free to compose agents with your own application code.
Installation
PydanticAI requires Python 3.9 or newer. Install it with pip:
pip install pydantic-ai
The base package pulls in support for the major model providers. If you want a slimmer install, you can select only the providers you need:
pip install "pydantic-ai-slim[openai]"
pip install "pydantic-ai-slim[anthropic]"
pip install "pydantic-ai-slim[google]"
If you plan to use observability, install the Logfire extra as well:
pip install "pydantic-ai[logfire]"
Most providers read their API key from an environment variable:
export OPENAIAPIKEY="sk-..."
export ANTHROPICAPIKEY="sk-ant-..."
export GEMINIAPIKEY="..."
Your First Agent
An Agent is the central object. You give it a model identifier and an optional system prompt, then run it.
from pydanticai import Agent
agent = Agent(
"openai:gpt-4o",
system
prompt="You are a concise assistant. Answer in one or two sentences.",
)
result = agent.runsync("What is the capital of Indonesia?")
print(result.output)
Jakarta is the capital of Indonesia.
The model string follows the pattern provider:model-name. Swapping providers is a one-line change:
agent = Agent("anthropic:claude-3-5-sonnet-latest")
agent = Agent("google-gla:gemini-1.5-flash")
You can also pass a model instance directly if you need custom configuration such as a base URL or timeout. The string form is convenient for the common case.
Three Ways to Run an Agent
PydanticAI offers three run methods depending on your execution context:
import asyncio
Synchronous - convenient in scripts and notebooks
result = agent.runsync("Hello")
print(result.output)
Asynchronous - for async applications and concurrency
async def main():
result = await agent.run("Hello")
print(result.output)
asyncio.run(main())
Streaming - receive output incrementally as it is generated
async def streamexample():
async with agent.runstream("Tell me a short story") as response:
async for chunk in response.streamtext(delta=True):
print(chunk, end="", flush=True)
asyncio.run(streamexample())
All three return (or yield) the same rich result object, which carries the output, the message history, and token usage.
Structured, Typed Outputs
Free-form text is fine for chat, but applications usually need structured data. PydanticAI lets you declare an outputtype as a Pydantic model. The framework instructs the model to produce matching data and validates the response before handing it back.
from pydantic import BaseModel, Field
from pydanticai import Agent
class CityInfo(BaseModel):
name: str
country: str
population: int = Field(description="Approximate population")
iscapital: bool
agent = Agent(
"openai:gpt-4o",
outputtype=CityInfo,
systemprompt="Extract structured information about the city the user names.",
)
result = agent.runsync("Tell me about Bandung")
print(result.output)
name='Bandung' country='Indonesia' population=2500000 iscapital=False
print(type(result.output))
main.CityInfo'>
result.output is a fully validated CityInfo instance, not a dictionary. If the model returns something that does not fit the schema, PydanticAI automatically asks it to correct the response, up to a configurable retry limit.
You can also use a union of types when an agent might return one of several shapes:
from typing import Union
class Success(BaseModel):
answer: str
class Clarification(BaseModel):
question: str
agent = Agent(
"openai:gpt-4o",
outputtype=Union[Success, Clarification],
)
The model picks the appropriate shape, and the result is typed as the union so you can branch on it safely.
Dependency Injection
Real agents need access to databases, HTTP clients, configuration, or the identity of the current user. Rather than reaching for globals, PydanticAI uses dependency injection. You declare a depstype, usually a dataclass, and pass an instance when you run the agent. Tools and system prompts then receive it through a RunContext.
from dataclasses import dataclass
import httpx
from pydanticai import Agent, RunContext
@dataclass
class SupportDeps:
customerid: int
db: "DatabaseConn"
httpclient: httpx.AsyncClient
agent = Agent(
"openai:gpt-4o",
depstype=SupportDeps,
systemprompt="You are a support agent for an online store.",
)
The depstype is generic, so RunContext[SupportDeps] gives you full autocompletion and type checking on ctx.deps. This makes testing far easier because you can inject fakes or stubs without monkeypatching.
Function Tools
Tools let the model call your code during a run. PydanticAI provides two decorators:
@agent.toolfor tools that need access to dependencies. The first parameter is aRunContext.@agent.toolplainfor tools that need no context.
PydanticAI builds the JSON schema for each tool from its type hints and docstring, so the model knows how and when to call it.
from pydanticai import Agent, RunContext
@agent.tool
async def get
customername(ctx: RunContext[SupportDeps]) -> str:
"""Return the current customer's full name."""
return await ctx.deps.db.customer
name(id=ctx.deps.customerid)
@agent.tool
async def get
balance(ctx: RunContext[SupportDeps], includepending: bool) -> float:
"""Return the customer's account balance.
Args:
include
pending: Whether to include pending transactions.
"""
return await ctx.deps.db.customerbalance(
id=ctx.deps.customerid,
includepending=includepending,
)
@agent.toolplain
def currentexchangerate(base: str, quote: str) -> float:
"""Return a static demo exchange rate between two currencies."""
rates = {("USD", "IDR"): 16250.0, ("EUR", "IDR"): 17600.0}
return rates.get((base, quote), 1.0)
The docstring description and the Args section are extracted automatically. Parameter types are validated when the model calls the tool, so a tool declared with includepending: bool will never receive a string.
Dynamic System Prompts
A static system prompt is set at agent creation. When the prompt depends on runtime data, register a dynamic one with @agent.systemprompt. It runs on every execution and can read dependencies.
@agent.systemprompt
async def addcustomercontext(ctx: RunContext[SupportDeps]) -> str:
name = await ctx.deps.db.customername(id=ctx.deps.customerid)
return f"The customer's name is {name}. Address them politely by name."
Static and dynamic system prompts are combined, with dynamic ones evaluated at run time. You can register several; they are concatenated in registration order.
Output Validators
Sometimes validating the shape of an output is not enough; you need to validate its content against business rules or an external system. The @agent.outputvalidator decorator runs after the model produces a structured output. If you raise ModelRetry, PydanticAI sends the error back to the model and asks it to try again.
from pydanticai import Agent, ModelRetry, RunContext
@agent.outputvalidator
async def validatesupportoutput(
ctx: RunContext[SupportDeps], output: "SupportResult"
) -> "SupportResult":
if output.blockcard and output.risk < 5:
raise ModelRetry(
"You set blockcard=True but risk is low. "
"Only block the card when risk is 5 or higher."
)
return output
This keeps your validation logic close to the agent and gives the model a chance to self-correct rather than failing outright.
Conversation History Across Runs
Each call to run, runsync, or runstream is stateless by default. To continue a conversation, pass the previous messages through messagehistory.
agent = Agent("openai:gpt-4o", systemprompt="You are a helpful tutor.")
first = agent.run
sync("My name is Ruby and I am learning Python.")
print(first.output)
second = agent.runsync(
"What was my name again?",
messagehistory=first.newmessages(),
)
print(second.output)
Your name is Ruby.
Use result.allmessages() to get the full history including the system prompt, or result.newmessages() for only the messages produced by the latest run. Messages can be serialized to JSON for storage between requests in a web application.
Streaming Structured Output
Streaming is not limited to text. You can stream a structured output and receive partial, progressively validated objects as the model generates them.
from pydantic import BaseModel
from pydanticai import Agent
class Report(BaseModel):
title: str
summary: str
bulletpoints: list[str]
agent = Agent("openai:gpt-4o", outputtype=Report)
async def streamreport():
async with agent.runstream("Summarize the quarterly sales data") as response:
async for partial in response.stream():
print(partial)
final = await response.getoutput()
print("Final:", final)
Each partial yielded is a Report validated in non-strict mode, which is useful for rendering progress in a UI before the full response arrives.
Testing Agents
Testing code that calls an LLM is hard if every test hits the network. PydanticAI ships two models for testing and a clean override mechanism.
TestModel calls your tools and returns plausible structured data without contacting a real model. FunctionModel lets you script exactly what the model does. Agent.override swaps the model within a context manager so your production code stays unchanged.
from pydanticai import models
from pydanticai.models.test import TestModel
Prevent accidental real requests during the whole test suite
models.ALLOWMODELREQUESTS = False
def testsupportagent():
deps = SupportDeps(customerid=1, db=FakeDB(), httpclient=FakeClient())
with agent.override(model=TestModel()):
result = agent.runsync("Is my card blocked?", deps=deps)
assert isinstance(result.output, SupportResult)
For precise control, FunctionModel receives the message history and the available tools, and you return the exact model response:
from pydanticai.messages import ModelResponse, TextPart
from pydantic
ai.models.function import FunctionModel, AgentInfo
def mymodellogic(messages, info: AgentInfo) -> ModelResponse:
return ModelResponse(parts=[TextPart("A scripted reply")])
def testwithfunctionmodel():
with agent.override(model=FunctionModel(mymodellogic)):
result = agent.runsync("Anything")
assert result.output == "A scripted reply"
This approach gives deterministic, fast, network-free tests of your agent logic, tools, and validators.
Observability with Logfire
Debugging an agent that makes several model calls and tool invocations per request benefits from tracing. PydanticAI integrates with Logfire, also from the Pydantic team, to capture each step.
import logfire
logfire.configure()
logfire.instrumentpydanticai()
All agent runs after this are traced automatically
result = agent.runsync("Hello")
Logfire records the prompts, model responses, tool calls and their arguments, retries, and token usage. Because instrumentation is based on OpenTelemetry, you can also export traces to other compatible backends if you prefer not to use Logfire's hosted service.
End-to-End Example: A Customer-Support Agent
The following example ties the concepts together. It defines a support agent with a typed output, a dependency carrying a database client and the current customer, two tools, a dynamic system prompt, and an output validator.
from dataclasses import dataclass
from pydantic import BaseModel, Field
from pydanticai import Agent, ModelRetry, RunContext
A small fake database to keep the example self-contained
class DatabaseConn:
names = {1: "Ruby Abdullah"}
balances = {1: 1250.0}
async def customername(self, id: int) -> str:
return self.names[id]
async def customerbalance(self, id: int, includepending: bool) -> float:
balance = self.balances[id]
return balance - 50.0 if includepending else balance
@dataclass
class SupportDeps:
customerid: int
db: DatabaseConn
class SupportResult(BaseModel):
advice: str = Field(description="Advice to give the customer")
blockcard: bool = Field(description="Whether to block the customer's card")
risk: int = Field(ge=0, le=10, description="Risk level from 0 to 10")
supportagent = Agent(
"openai:gpt-4o",
depstype=SupportDeps,
outputtype=SupportResult,
systemprompt=(
"You are a support agent for a bank. Help the customer with their query "
"and judge the risk level of their request."
),
)
@supportagent.systemprompt
async def addcustomername(ctx: RunContext[SupportDeps]) -> str:
name = await ctx.deps.db.customername(id=ctx.deps.customerid)
return f"The customer's name is {name!r}."
@supportagent.tool
async def customerbalance(
ctx: RunContext[SupportDeps], includepending: bool
) -> str:
"""Return the customer's current account balance.
Args:
includepending: Whether to include pending transactions in the figure.
"""
balance = await ctx.deps.db.customerbalance(
id=ctx.deps.customerid,
includepending=includepending,
)
return f"${balance:.2f}"
@supportagent.outputvalidator
async def validaterisk(
ctx: RunContext[SupportDeps], output: SupportResult
) -> SupportResult:
if output.blockcard and output.risk < 5:
raise ModelRetry(
"Do not block the card unless the risk level is 5 or higher."
)
return output
async def main():
deps = SupportDeps(customerid=1, db=DatabaseConn())
result = await supportagent.run("What is my balance?", deps=deps)
print(result.output)
# advice='Your current balance is $1250.00.' blockcard=False risk=1
followup = await supportagent.run(
"I think my card was stolen, please help.",
deps=deps,
messagehistory=result.newmessages(),
)
print(followup.output)
# advice='I will block your card immediately...' blockcard=True risk=8
import asyncio
asyncio.run(main())
This single agent demonstrates the full picture: dependency injection supplies the database and customer, the dynamic system prompt personalizes responses, the tool fetches live data, the output validator enforces a business rule, and messagehistory keeps the conversation coherent across runs.
Best Practices
- Let the type system work for you. Always annotate
depstypeandoutputtype. The static guarantees catch mistakes before runtime and improve editor support. - Prefer dependency injection over globals. Passing a dataclass of dependencies keeps agents testable and makes data flow explicit.
- Write clear tool docstrings. The model relies on the docstring and type hints to decide when and how to call a tool. Treat them as part of the prompt.
- Use output validators for business rules. Schema validation checks shape; output validators check meaning. Raising
ModelRetrylets the model recover gracefully. - Set sensible retry limits. Automatic retries are helpful but can multiply cost. Configure
retrieson the agent or per tool to bound them. - Test with
TestModelandFunctionModel. SetALLOWMODELREQUESTS = Falsein your test suite to guarantee no accidental network calls. - Instrument from the start. Adding Logfire early makes diagnosing odd model behavior far easier than retrofitting logging later.
- Keep agents focused. Compose several small, single-purpose agents rather than building one monolithic agent with dozens of tools.
Conclusion and Key Takeaways
PydanticAI brings a disciplined, type-first approach to building LLM agents. By leaning on Pydantic for validation and adopting FastAPI-style dependency injection, it lets you build agents that are predictable, testable, and easy to reason about.
Key takeaways:
- An
Agentcombines a model, a system prompt, a dependency type, and an output type into one typed object. - Structured outputs with
outputtypegive you validated Pydantic models instead of raw text. - Dependency injection via
depstypeandRunContextkeeps agents clean and testable. - Tools (
@agent.tool,@agent.toolplain), dynamic system prompts, and output validators (ModelRetry) cover the common patterns of real applications. messagehistory, streaming,TestModel/FunctionModel, and Logfire round out the path from prototype to production.
Start with a single agent and a typed output, then add tools and dependencies as your application grows. The framework rewards a gradual, type-guided approach.