Accuracy is the wrong number to pick a model on

Two models that score the same on a benchmark can cost ten times apart per correct answer. Here is the number a finance team should look at instead.

10×apart in cost per correct answer, at the same accuracy

When a team picks a model for a task, the first number they ask for is accuracy. It is the wrong first number.

Accuracy tells you how often the model is right. It does not tell you what a right answer costs, how often the model guesses when it should have stopped, or whether a cheaper model would have done the same job.

Cost per correct answer

Take the same task, invoice matching, and run it across several models on a month of real cases. Score each one on three things:

  1. How often it matched correctly.
  2. How often it flagged an invoice it was unsure about, instead of guessing.
  3. What each correct match cost, in tokens and in the human time spent on the ones it got wrong.

Divide the total cost by the number of correct answers. That is the number to compare. In our benchmarks, models that sit within two points of each other on accuracy can be far apart on this number, because one of them guesses and the other flags.

Why flagging matters more than accuracy

A wrong match that goes through quietly costs more than a flag. Someone finds it later, or nobody does.

A model that says "I'm not sure, a person should look" on the hard two percent is doing the job right.

That holds even if a benchmark counts it as a miss. A controller would rather clear forty flags a month than find one silent error in the ledger.

In practiceSet the flag threshold with the team that clears the flags. They know what a false alarm costs them; the model doesn't.

What to do with the number

Once you have cost per correct answer per task, model choice stops being a debate. Route each task to the model that wins on that task. When a cheaper model catches up, move, without rebuilding the agent.

Our cost-per-correct-answer paper has the full method and the results across the models we test. Read it here.

More on the blog

From the research

Bring us a process your team still does by hand.

We’ll give you automated workflow agents.