A clause classifier, and an account of where it fails.
Reaching for a frontier model is the reasonable default, because it is the thing that works on everything. On one narrow, high-volume task it is the expensive way to be less accurate.
That is the measurement, not the argument. Everything below is how it was established — including the tier that turned out to be unnecessary, and the case where the model is confidently wrong and says nothing.
One clause in. A label, whether the model will stand behind it, and the shape of its certainty across all 100 classes.
No clause loaded. The instrument is on and waiting.
The same weights. The same 3,000 rows. One CPU instruction apart.
The tier everyone assumes you need.
The flag catches text that looks different. It does not catch text that looks like a contract and is not one.