Model benchmarks

Contract review: where fourteen models hold up, and where they don't

· Benchmark

Fourteen models, one set of real commercial contracts, one question: which of them can be trusted with a first-pass review, and on which clause types.

The benchmark scores each model on clause extraction, risk scoring against a standard, and the false-negative rate that matters most to counsel: the risky clause it did not flag.

What we tested

Two hundred commercial agreements, anonymised, across MSAs, NDAs, order forms, and vendor terms. Each model extracted the clauses counsel cares about, scored each against the company's standard, and said which ones a person should read.

Every model ran with the same instructions and the same standard. We measured extraction recall, scoring agreement with two lawyers, and the rate at which a model missed a clause the lawyers flagged.

What we found

Extraction is largely solved: eleven of fourteen models found more than ninety-five percent of the clauses. Scoring is not. Agreement with counsel ranged widely, and the models that scored best on easy clauses were not the ones that missed fewest risky ones.

The number that should decide the choice is the missed-risk rate. On that number, three models were in a class of their own, and one of them was among the cheapest to run.

Model AModel BModel CModel D
Figure 1. Missed-risk rate by clause type across the fourteen models. Lower is better. Placeholder data.

What it means for your team

A first-pass review agent is ready for production on the right model, with a person on every clause it scores above the threshold. It is not ready on the wrong one, however good its demo looks.

The paper names the models and the numbers. The PDF opens with a work email.

Read the paper

Enter a work email and the PDF opens. We'll send you the next paper when it's out, and nothing else.

Get the next paper when it's out.

One email per paper. Nothing else.