Accuracy alone doesn't tell a finance team which model to run. This paper measures what a correct answer costs, per task, across the models in our benchmarks.
It also shows when a cheaper model catches up, and how to switch without rebuilding the agent.
What we tested
We took a month of real invoice matching and contract review cases from a mid-sized company, removed anything identifying, and ran each case through every model under the same instructions and the same tools.
Each model saw the same context, in the same order, with the same budget of steps. Where a model asked for a person, we counted that as a flag, not a failure.
What we found
The models split into two groups. The first got the easy cases right and guessed on the hard ones. The second got slightly fewer easy cases right and flagged the hard ones. For a finance or legal team, the second group is the one to trust.
Cost per correct answer, not accuracy, separated the models most. Two models within two points of each other on accuracy were a factor of six apart on what a correct answer cost.
What it means for your team
Pick the model per task, not per company. Route the easy cases to the cheapest model that clears the bar and keep the expensive one for the cases that need it.
Test on your own cases before you trust any table, including this one. The benchmark is a starting point. The regression suite on your data is the thing that keeps it true.