Reasonable Doubt

A clause classifier, and an account of where it fails.


the same 3,000 contract clauses
a frontier model, zero-shot
—
—
against
an encoder fine-tuned for the job, on a CPU
—
—

Reaching for a frontier model is the reasonable default, because it is the thing that works on everything. On one narrow, high-volume task it is the expensive way to be less accurate.

That is the measurement, not the argument.   Everything below is how it was established — including the tier that turned out to be unnecessary, and the case where the model is confidently wrong and says nothing.


01

Try it

One clause in. A label, whether the model will stand behind it, and the shape of its certainty across all 100 classes.

statuschecking…

 

Paste a clause, or try
0 / 20,000

—

idle

No clause loaded. The instrument is on and waiting.

02

Cost

 

—
cheaper per clause
served here
—
the API
—
per month
—

 

03

Precision

The same weights. The same 3,000 rows. One CPU instruction apart.

—
FP32 — macro-F1
—
INT8 — chance, on 100 classes

 

04

Escalation

The tier everyone assumes you need.

—
 

 

05

Where it fails

The flag catches text that looks different. It does not catch text that looks like a contract and is not one.

—
 
—