Note #39 •

Your Coding Agent Is Overpaying for Intelligence It Never Uses

A coding agent that always uses the frontier model is paying for insurance it never claims

LangChain's Open SWE is an open source coding agent that engineers drive from Slack and a web UI. Before building a router, the team pulled a week of interactive threads out of its traces and labelled each one by task type. New features were 22% of threads, bug fixes 17%, and test or no-op runs 16%.

Every one of those threads was running on a top-tier frontier model. Feature investigations were long and expensive. Test and release runs were short and cheap. The mix suggested the obvious hypothesis. Most requests don't need frontier intelligence, and a router could read the incoming request and pick a cheaper model without degrading the outcome.

The numbers

The first A/B test split 973 threads. Half went through the router, half always used the strongest model.

Quality did not move. 29.2% of routed threads ended in a merged PR against 27.3% for the control, p = 0.49. PR open rates were 38.9% against 39.6%, p = 0.82.

Cost moved a lot. The median routed thread cost $0.94 against $2.61, a 64% drop. The mean fell 42% and the 90th percentile fell 37%, so the saving was not a handful of cheap outliers dragging an average around.

The router sent 56% of threads to the balanced tier, 34% to fast, and only 10% to performance. The ladder between tiers is steep. Median thread cost was $0.097 on fast, $1.50 on balanced, and $2.88 on performance. That is a 30x spread between the cheapest and the most expensive choice.

Why the router belongs in the harness

A gateway sees a request. A harness sees the task.

That distinction carries the whole argument. The criteria for each tier were written from Open SWE's own task breakdown, not from a general benchmark. The base prompt and the three tier descriptions are specific to the work Open SWE handles, which is investigations, code changes, review, and release procedures. A generic gateway cannot write those criteria, because it does not hold the agent's prompt, tools, or domain knowledge.

In LangChain, the decision point is middleware. Middleware can swap the model an agent calls without changing anything else about the agent, so routing is a change to one file rather than a rewrite of the agent.

The classifier is the cheap part

The router runs on the thread's first human message. It has three parts. A base prompt tells the classifier its job, which is to pick the least expensive model likely to complete the task. Criteria per tier describe in plain language what work each tier should take. A classifier model reads the request and picks a tier.

Two implementation details are worth stealing.

The first version used an LLM with structured output. It now runs on Jev, a decision model from TypeSafe, which made classification almost 50x faster.

The second is in the code. The router reads only the latest human-authored message and truncates it to the last 8,000 characters. That filter matters more than it looks. Context blocks get appended to the message list as human messages, so the newest message is usually machine-authored. Taking the newest message without checking its kind would route on injected context instead of on the request.

The experiment that failed in a day

The team also ran the opposite test. Router against always using the fast model.

They killed it within a day, before it could produce statistically meaningful results. Engineers flagged the fast-only arm almost immediately and said the low output quality was disrupting their work.

That failure is the more useful result. Routing everything to a cheap model is not a cost optimization, it is an outage. The asymmetry is the point. A task running on a model too weak for it does not cost less. It costs a redo, and the redo is where the real spend hides.

The bill that arrives when you switch models mid-thread

The router picks once, at the start of a thread, and that model handles the whole thread. Re-routing mid-thread sits in the future work section, and the reason is the prompt cache. Switching models throws it away, so the new model re-reads the entire thread at full price.

For async agents that cost is often negligible, because the cache expires between human turns anyway when the time to live is short, like five minutes. For a synchronous agent where turns arrive seconds apart, a mid-thread switch can cost more than it saves. Anyone reaching for mid-thread routing should price the cache first.

What to do before you route anything

The order in the post is the order that matters.

Understand the tasks first. Your traces hold the record of what the agent is actually asked to do, and you cannot pick tiers until you know the mix. Cost and invocation count are the cheap complexity signals. The post uses both, and notes that a high invocation count can mean either a hard task or a model that needed follow-ups.

Then pick models along the cost-intelligence curve. The Pareto frontier is the set of models that are cheapest and smartest, and three points on it are enough to start.

Then build the router in the harness and treat it as context engineering. Decide what the router needs to see to judge model fit for your domain.

Then, and this is the step teams skip, put outcome measurement in place before you route. Open SWE records every PR it opens and whether it merged, and merged PRs per thread became the main success metric. Thumbs up and down covered the threads that never produce a PR. Offline evals are cleaner to run but need a dataset that looks like real traffic and is graded on what users care about, which for a coding agent means PR quality and reviewability. That is hard to grade offline, which is why the live A/B test carried the argument.

A router that cuts cost and cannot show its quality is flat is not a result. It is an unmeasured change to production.