Why publish the terms first
A benchmark is only worth reading if you can check it, and the way to prove that is to write down the rules before there is a result to be embarrassed by. What follows is binding on us.The terms
The methodology is published
The methodology is published
How each task is posed, how an answer is scored, and what counts as a
failure — written down and public, before any score is.
The harness is reproducible
The harness is reproducible
Someone else runs it and gets our numbers, or the numbers do not stand.
Every number comes from an actual run
Every number comes from an actual run
No estimates, no figures carried over from a vendor’s own page, no cells
filled in by inference.
Only models we can actually call are scored
Only models we can actually call are scored
A model we cannot run is listed as not yet run. It never gets a score,
and a competitor’s product never gets one we invented for it.
The report states its limits
The report states its limits
What the tasks measure is not “which legal AI is best”. The report says what
it measured and what it did not, in its own words, at the top.
What it will measure
The task set is being designed and is not published yet. It will cover the work this product is built to do with contracts — reading a clause, proposing a change, and saying plainly what a document says — and the definitions will land here in full before the first run, along with the report. Until then there is nothing on this page to cite, which is the point.What Lex is designed to do
The job, and the two things it will never do.