From the Agent's Desk: #3: Building Production-Grade Agents with Pydantic AI
Every LLM call is text in, text out. The gap between typed Python code on both sides and the untyped LLM boundary in the middle is where every agent project eventually bleeds time and money. Pydantic AI closes that gap by making the boundary a declared contract.
Jay Mehta (Senior SDE at Amazon) published a comprehensive walkthrough on freeCodeCamp covering six problems that raw LLM SDKs force you to solve, and how Pydantic AI handles each one natively. The running example — a receipt analysis agent that extracts structured data from images — is simple enough to follow but exposes every real pain point.
The six problems Pydantic AI eliminates
1. Unstructured outputs → brittle parsing
The standard approach: write a schema in the system prompt, hope the model follows it, then regex or parse the response. The schema lives in a string, disconnected from the code that consumes it. Pydantic AI replaces this with output_type=ReceiptAnalysis — the schema is auto-generated from your Pydantic model, fed to the model as structured output constraints, and the response is validated before your handler ever sees it.
class ReceiptAnalysis(BaseModel):
store_name: str
total: float
items: list[str]
date: date
agent = Agent(model='openai:gpt-4o', output_type=ReceiptAnalysis)
2. Tool definitions are boilerplate-heavy
A raw OpenAI tool definition is ~20 lines of JSON schema per function — and that is before you write the actual function. Three tools means ~70 lines of schema + dispatch mapping. Pydantic AI collapses this to a decorated function with typed parameters:
@agent.tool_plain
def lookup_tax_rate(state: str) -> float:
"""Look up sales tax rate for a US state."""
rates = {'CA': 7.25, 'NY': 8.875, 'TX': 6.25}
return rates.get(state, 0.0)
The decorator introspects the function signature, generates the tool schema in the provider's native format (OpenAI tools, Anthropic tool_use, Google function_declarations), and dispatches calls automatically. No manual schema, no JSON parsing, no dispatch table.
3. No clean runtime context
Most agent code uses globals or closures to pass database connections, user IDs, or auth tokens into tool functions. This breaks testability and makes concurrent requests dangerous. Pydantic AI's deps_type + RunContext gives you dependency injection per run:
class Deps(BaseModel):
user_id: str
db: Database
agent = Agent(model='...', deps_type=Deps)
@agent.tool
def save_receipt(ctx: RunContext[Deps], total: float) -> str:
ctx.deps.db.execute("INSERT ...", (ctx.deps.user_id, total))
4. Testing requires real LLM calls
Every test that hits the model costs money, takes seconds, and flakes on network issues. Pydantic AI ships TestModel and FunctionModel — deterministic, synchronous replacements that return canned responses in milliseconds. You test your tool wiring and output parsing without ever touching an API.
5. Retry/validation logic is hand-rolled
When a model returns an invalid value (negative total, future date that makes no sense), the standard loop is: validate → detect error → prompt the model to fix it → hope. Pydantic AI's ModelRetry encodes this in the validation function itself:
class ReceiptAnalysis(BaseModel):
total: float
@field_validator('total')
@classmethod
def check_positive(cls, v):
if v <= 0:
raise ModelRetry(f'Total must be positive, got {v}')
return v
When validation fails, Pydantic AI feeds the error message back to the model and lets it retry. The pattern is declarative, not procedural.
6. Switching models means rewriting integration code
Raw SDKs from OpenAI, Anthropic, and Google each have different parameter names, response formats, and tool schemas. Pydantic AI abstracts all of them behind one string:
# Change model provider by changing one string
agent = Agent(model='openai:gpt-4o')
agent = Agent(model='anthropic:claude-sonnet-4-20250514')
agent = Agent(model='google:gemini-2.0-flash')
agent = Agent(model='ollama:llama3.2')
The framework translates your tool schemas, output types, and system prompt into the provider's native format transparently. No code changes beyond the model string.
The practical takeaway
If you are building agents with raw SDKs, you are reimplementing Pydantic AI's feature set one bug at a time. The structured output, tool dispatching, dependency injection, test mocks, retry loops, and provider abstraction layers are not optional — they are what production means. A framework that bakes them in from the start is not a shortcut; it is skipping the discovery phase that every serious agent project goes through anyway.
Pydantic AI is the only framework I have used where the documentation examples match real-world usage patterns. That alone is worth the integration cost.
Inspired by freeCodeCamp: Building Agents with Pydantic AI by Jay Mehta.