Model benchmarks

The monthly close, model versus model

· Benchmark

Invoice matching, accrual drafting, and reconciliation checks, run model against model on a month of real finance data.

Scored on accuracy, on how often a model guesses where it should flag, and on the cost of a correct answer.

What we tested

We took a month of real month-end close cases from a mid-sized company, removed anything identifying, and ran each case through every model under the same instructions and the same tools.

Each model saw the same context, in the same order, with the same budget of steps. Where a model asked for a person, we counted that as a flag, not a failure.

What we found

The models split into two groups. The first got the easy cases right and guessed on the hard ones. The second got slightly fewer easy cases right and flagged the hard ones. For a finance or legal team, the second group is the one to trust.

Cost per correct answer, not accuracy, separated the models most. Two models within two points of each other on accuracy were a factor of six apart on what a correct answer cost.

Model AModel BModel CModel D
Figure 1. Matching accuracy against flag rate, per model, on one month of close data. Placeholder data.

What it means for your team

Pick the model per task, not per company. Route the easy cases to the cheapest model that clears the bar and keep the expensive one for the cases that need it.

Test on your own cases before you trust any table, including this one. The benchmark is a starting point. The regression suite on your data is the thing that keeps it true.

Read the paper

Enter a work email and the PDF opens. We'll send you the next paper when it's out, and nothing else.

Get the next paper when it's out.

One email per paper. Nothing else.