Routing coding agents is harder than it looks
A small experiment on whether routing in interactive sessions can save you a dime.
2026-10-01 · Arseny Kravchenko

Every few weeks someone pitches model routing for coding agents as an easy win: look at the prompt, send the simple stuff to a cheaper model, summon the frontier one for the hard problems. OpenRouter shipped their Auto Router, a lightweight classifier that puts each prompt into one of ~30 task types and picks a model. Databricks benchmarked coding agents on their own codebase and concluded they "should push more work to the Haiku and GPT 5.4 Mini class of models."
The conclusion was too tempting to ignore, but my motivation was less noble: Fable limits run out, Opus limits mostly don't. So, I spent a weekend checking whether my own history supports the pitch.
TL;DR: it doesn't. At least not in the easy form, and the reasons are more interesting than "classifiers are bad". (The weekend itself ran mostly on Fable. Guess where the limits went.)
Anecdotal: one person's bill, two harnesses, and a judge that is itself a model. Take it with a grain of salt.
The setup
The corpus is every local session I had: 3,135 in total, 214 from Claude Code and 2,921 from Codex. Most of the Codex sessions were launched by Claude Code as helpers. I started 329 of them myself. At API list prices the whole pile is about $10.9k, and those 329 I started are $7.5k of it.
There was one question for every session: does this task require a frontier model rather than a mid-tier one? I asked it twice, using Jev (TypeSafe's small judgment model, picked mostly for fun):
-
p_first: from the opening message only. This is what a router would see.
-
p_full: from the whole session, compressed to ~4% of its context (user turns, assistant prose, tool calls with truncated arguments, no tool output). This is the retrospective verdict.
These tasks were run only once, so there is no ground truth. Everything below is agreement between two views of the same session and a few sanity checks against other readers.
The Yes: The signal is there
On 98 clean sessions (fresh openings, no injected preambles, no sub-threads), p_first predicts the retrospective verdict at AUC 0.79 (0.5 is a coin flip, 1.0 is a perfect ranking of hard above easy; the 95% interval is 0.70–0.88). Not bad for one message.
However, there are two caveats that deflate it:
-
The retrospective judge reads length. A long transcript looks like a hard task. The first time around this made "difficulty is mostly scope" look like the finding: the opening plus the realized session length hit AUC 0.88. Then I built a length-equalized verdict (the same judge read the first and last 2,000 characters only) and had three Opus readers label 30 sessions with an instruction to ignore trace length. The readers agreed with each other almost perfectly, and they sided with the equalized judge far more than with the original one (Cohen's kappa, an agreement score where 0 is chance and 1 is perfect, hit 0.86 vs 0.57). Against them, the realized length predicts no better than the opening, while the opening itself drops to 0.68. The share of "hard" sessions falls from 69% to 50%.
-
Terse prompts hide the work. Out of those 98 clean sessions, the opening underrates difficulty in 86. A one-line request for a short cartoon with three characters and an initial score of 0.12 grew into 241 turns and got 0.78 in hindsight. A bare paper link plus "What should I take from here?" went from 0.27 to 0.82. The real difficulty lives in the repo and in what I have in my head, but not in the text. A real router would at least see the repository; mine didn't.
Still, a moderate signal exists, and a router would be a reasonable weekend project if the story ended here. Courtesy of my motivation, it doesn't.
The But: The signal doesn’t count the money
Spend on my sessions follows a steep power law:
-
Sessions over 200k tokens: 31% of sessions, 91% of the cost.
-
Sessions under 50k: 24% of sessions, 0.2% of the cost.

The awkward part: all 98 labelled sessions are under 200k tokens, because longer ones didn't fit the pipeline (skill issue, admittedly). So the classifier above was validated on the 9% of the bill where the decision is least important, and the tail's hard rate is an assumption. The top labelled bin (160–200k) is 75–100% hard depending on the verdict. Below I assume a more forgiving 64% or 80%: at 100%, routing loses at every threshold.
Counting by sessions, a router looks great: send the bottom ~20–30% cheap and save 15–28% at the most generous price ratio. Weighting by dollars is where it falls apart.
What the tail looks like
The judge didn’t see the tail, but I can read it. There are 101 sessions over 200k. The median session has 13 user messages and runs for 12 hours of wall clock, with 92% of the sessions spanning at least five messages. So these are not some fire-and-forget tasks. We’re talking about long, drifting conversations. The top 15 most expensive sessions totaled around $3,900 (about half of the $7.5k I started), with a few notable examples:
-
A one-liner request, “draft a list of integrations to build next,” became a plugin marketplace redesign, a 30k-line PR review, and a full end-to-end run. It took 100 messages over nine days and cost $665, the single most expensive session I have.
-
"Use .env for openrouter key and rerun my benchmarks … write a detailed memo." A textbook cheap task (run a script, summarise) frankensteined into five days of new comparison arms, scorer audits and 3,900-sample matrices across three models, for $269.
-
A pasted error accompanied by "Why didn't this work?" ended up redesigning the core semantics of how unknown tools are treated. That took 26 messages of me arguing with the agent about lattices and 12 review rounds. $231.
-
A prompt asking which simple (taste the irony) components to ship next triggered a schema change in the core engine and cured me of using the word ‘simple’ for some time.
About half of the top 15 start with a long, pre-written spec handed to a skill I use for autonomous builds. The skill plans, gets the plan reviewed, then implements for hours with review gates along the way. These kinds of sessions announce their difficulty, and a router would correctly keep them on the frontier model. The other half look deceptively cheap or moderate from the first message: a question, a link, "draft a list". Those are exactly the ones that the turn-0 router sends one tier down—and pays for twice. None of the 15 is cheap as a whole. Almost every one has cheap parts: CI fixes, rebases, status checks, PR description rewrites. Keep that in mind for the subagent section.
There are counterexamples, and the best one sits lower in the tail. A 94-message session running a message campaign for a side project with a final cost of $76. It included copy rewrites, dry runs and "check status" requests for three weeks, with two genuinely tricky bugs in the middle. A mid-tier model would have been fine for 90% of it, but a router would have had no way to tell from the opening.
A botched routing decision costs more than a correct one saves
A cheap run costs c of a frontier run. A hard task sent cheap costs c + 1 + h:
-
c for the failed attempt.
-
1 for the redo on the frontier model.
-
h for my time and cleanup, in units of one frontier session.
At h = 0, routing only pays if fewer than 1 − c of the cheap bucket's cost turns out hard.
The pitch-deck math assumes c ≈ 0.2, but real adjacent tiers are closer: Sonnet 5 is 0.4 of Opus 5, and Opus 5 is 0.5 of Fable. Here’s how savings at h = 0 at the best threshold look, weighted by cost across the sessions I started (the tail assumed 64% or 80% hard):
| price ratio c | tail 64% hard | tail 80% hard |
|---|---|---|
| 0.2 | +7.6% to +9.2%, break-even h ≈ 0.25 | +0.3% to +1.5% |
| 0.4 | loses | loses |
| 0.5 | loses | loses |
| each Claude session one tier down | loses | loses |
"Loses" here means negative at every threshold before counting any human time, typically −0.2% to −4%. At c = 0.2 the router saves single digits and stops paying once a botched task costs more than about a quarter of a frontier session in cleanup. A confidently executed wrong refactor I have to untangle clears that bar easily.

The list price overstates the gap for another reason: caching. Coding agents live on cached input. Fable 5.1 reads cache at $0.25/MTok and Opus 5 at $0.50/MTok: the cheaper model has the pricier cache reads. So on a long, cache-heavy session, dropping from Fable 5.1 to Opus 5 costs a median 0.83 of the original, not 0.5. Databricks saw the per-task version of this: Sonnet 5 cost "$2.09/task vs Opus's $1.94", "consuming 1.9x more tokens". Cheaper per token isn't cheaper per task.
Databricks' own conclusion doesn't contradict any of this. They tagged "about a quarter" of their logged usage as low-complexity, which is roughly my per-session cheap bucket, too. I don't think the difference is the workload. The difference is in the measurement unit:
Per task, a quarter of the work is cheap.
Per dollar on an interactive agent, a quarter of the sessions is a rounding error.
What about waiting one turn?
The obvious fix would be to route after the second user message instead of the first. On the 78 sessions with at least two user messages, AUC goes 0.78 → 0.92, which looked like a big win until I split by how much of the session the two-message prefix already covers:
-
Prefix covers ≥50% of the transcript: 0.72 → 1.00. The judge is grading the answer key.
-
Prefix covers <50%: 0.82 → 0.87, +0.05 [−0.05, +0.16]. Not significant.
Meanwhile, 17–37% of the cost is already spent at this point, and switching models brings us back to the caching issue. Microsoft's Copilot study measured an 8% cache hit rate after a model switch vs 55% across same-model turn boundaries. Anthropic says it bluntly: "It would actually be more expensive to switch to Haiku than to have Opus answer, because we would need to rebuild the prompt cache for Haiku".
The switch cost is asymmetric, though. Rewriting the median 74k context costs about 8% of a session when you upgrade, and 1.6–4% when you downgrade. Keep that asymmetry in mind for later.
Can subagents help?
Sessions aren't uniformly hard. Inside a long session, there are stretches of searching, diagnosing, and bookkeeping, and their details can be dropped once there's an answer. Hand such a stretch to a cheaper subagent and you save twice: a lower-tier model is used for the stretch, and the parent won’t carry those tool outputs in its context ever again.
It turns out I already route, just blindly. In 274 recent spawns, 114 ran a tier below their parent (97 of them Fable → Opus), because the tier is fixed by the agent type, not by the task. Spawning a subagent is cheap (median boot is $0.095, about 5% of a spawn's cost), and over-delegation is negligible ($1.52 in total on trivial spawns).
To count what nobody delegated but could have, I needed a closure test. A stretch is closed if the files and search targets it touched don't come back later in the session. My first version was wrong in an instructive way. An explore pass walks half the repo to name the files for the main agent to fix. That's a textbook delegation, but it counts as a leak. After several rounds of fixes, the metric agrees with real delegations: spawned work leaks less than matched stretches of the same session in 71% of pairs.
The supply of delegable stretches nobody delegated is ~5% of parent spend (154 stretches, $219 on a $4.5k parent bill). About two-thirds of that is context carry, not the tier discount. 19 of the 154 would cost more one tier down, thanks to that cache-read pricing again.
Then the part that highlights a flawed approach to cheap delegation. When I asked readers which stretches could be handed off on a one-sentence report and asked Jev which stretches were "mechanical rather than judgment," the two were strongly inverted: AUC 0.11 (well below the coin-flip 0.5, so the ranking runs backwards; 95% interval 0.00–0.24). That's 36 items, so treat it as a lead, not a result. What can be closed off cleanly is diagnosis that ends in a located cause, which is judgment work. The mechanical stuff, like editing a run of related files, is exactly what the parent needs in its context afterwards. The work that is easiest to hand off may be the work a weaker model is least suited for. Whether a weaker model actually copes I couldn't observe at all. A gate like
cargo test isn't an oracle when deleting the test also makes it pass.Can a hook do it automatically?
I simulated two hooks offline:
| hook | ceiling (perfect classifier) | best real rule | false nudges / session |
|---|---|---|---|
| on user prompt | 1.7% | 0.5% | 1.75 |
| on tool call (read-only streak touching new files) | 3.9% | 1.4% | 3.2 |
The prompt hook is dead on arrival. About two-thirds of the value follows no user prompt at all: it comes after compaction, task notifications or slash commands. The prompts that do precede it are things like go, applied, and let's rebase. The tool hook works, weakly. Held to at most one false nudge per session, it saves 0.4%. Mid-stretch, a search that will end in a finding looks exactly like reading code before editing it.
So instead of a hook I went with a very spartan solution: one paragraph in my global CLAUDE.md. The most useful output of the whole exercise turned out to be not a number but the criteria for what (and what not) to delegate:
-
A stretch of several tool calls where only the conclusion is relevant afterwards: locating code, diagnosing a failure from logs, probing the environment, gathering stats.
-
Not the core change of the task and not reading code the parent will build on next.
-
Sent to a subagent one tier below the parent (today that's Opus under Fable; by the time you read this it may be Sonnet under Opus), with a self-contained brief and a clear success condition.
-
With meeting the condition by weakening the check (skipping tests, loosening thresholds) explicitly ruled out.
Minutes later the agent cited it to hand off a review. That's the entire evaluation so far: a vibe check, not a benchmark. I haven't measured it; rerunning the simulation on post-change sessions is on the list.
The verdict
A text-only turn-0 router on an interactive, prototyping-heavy workload saves low single digits at a price ratio real tiers don't offer, and loses money at the ones they do. It's not just my setup: Agent-as-a-Router calls this the "information deficit" of static routers, and Bai et al. found that runs of the same task differ by up to 30× in tokens and that models can't predict their own spend.
What I'd try next (untested): start on the frontier model and downgrade late, since that's the cheap direction for the cache; route on the repo (diff size, test coverage, touched modules) rather than the prompt; delegate for context, not for price. And none of it counts until the same openings are replayed on both tiers with the tests frozen, turning a model agreeing with itself into an actual outcome.
Until then: teach the strong model to delegate rather than make the weak one guess what the task is.
