Skip to main content
ClauseBench is in preparation. There is no methodology document, no dataset, no harness, no run and no leaderboard yet — so there are no numbers on this page, and there will be none until there is a run behind them.

Why publish the terms first

A benchmark is only worth reading if you can check it, and the way to prove that is to write down the rules before there is a result to be embarrassed by. What follows is binding on us.

The terms

How each task is posed, how an answer is scored, and what counts as a failure — written down and public, before any score is.
We write the material and we say so. Where it comes from, how it was made, and what it is not representative of.
Someone else runs it and gets our numbers, or the numbers do not stand.
No estimates, no figures carried over from a vendor’s own page, no cells filled in by inference.
A model we cannot run is listed as not yet run. It never gets a score, and a competitor’s product never gets one we invented for it.
What the tasks measure is not “which legal AI is best”. The report says what it measured and what it did not, in its own words, at the top.

What it will measure

The task set is being designed and is not published yet. It will cover the work this product is built to do with contracts — reading a clause, proposing a change, and saying plainly what a document says — and the definitions will land here in full before the first run, along with the report. Until then there is nothing on this page to cite, which is the point.

What Lex is designed to do

The job, and the two things it will never do.