We have published our contract review benchmark. It runs fourteen models against one set of real commercial contracts and asks a single question: which of them can be trusted with a first-pass review, and on which clause types.
Each model is scored on clause extraction, on risk scoring against a standard, and on the number that matters most to counsel: the risky clause it did not flag.
The paper is on the research page. A work email opens the PDF.