Across a year of benchmarks, the cheapest model that clears a task's bar changed four times. This paper shows how to move each time without touching the agent.
Placeholder summary. The real paper's findings replace this paragraph.
What we tested
We took a month of real model routing cases from a mid-sized company, removed anything identifying, and ran each case through every model under the same instructions and the same tools.
Each model saw the same context, in the same order, with the same budget of steps. Where a model asked for a person, we counted that as a flag, not a failure.
What we found
The models split into two groups. The first got the easy cases right and guessed on the hard ones. The second got slightly fewer easy cases right and flagged the hard ones. For a finance or legal team, the second group is the one to trust.
Cost per correct answer, not accuracy, separated the models most. Two models within two points of each other on accuracy were a factor of six apart on what a correct answer cost.
What it means for your team
Pick the model per task, not per company. Route the easy cases to the cheapest model that clears the bar and keep the expensive one for the cases that need it.
Test on your own cases before you trust any table, including this one. The benchmark is a starting point. The regression suite on your data is the thing that keeps it true.