How we think about numbers.
The rubrics are private. The reasoning behind them is not.
Six rules, written before the work starts.
Kahneman, Sibony and Sunstein call the remedy for noisy judgment decision hygiene. These six rules are our version of it.
Write the rubric before you see the evidence.
Adjust it mid-analysis and you fit conclusions to findings, and you lose the ability to compare anything to anything. The sciences call this pre-registration. It works here for the same reason.
Every score gets a written anchor.
A 0 to 5 with no anchor is astrology with decimals.
Declare what you don’t know, then compute what it could change.
A verdict that cannot say “no answer changes this” is not a verdict yet. It is a guess with a number on it.
Show the arithmetic on every number.
A test that’s 99% accurate, looking for something that happens 0.1% of the time, is wrong about 90 times out of 91. That is not a flaw in the test. It is what the base rate does to it.
A process that greenlights everything is a random number generator.
The kill rate is the product. A screen that never says no has measured nothing.
Test the instrument the way you would test the judge.
Score the same case written three ways, facts held constant and only the prose changed, and watch how far the number moves. If style moves it, the rubric is scoring writing rather than substance.
One gate, one question, one number.
Work runs as an investigation, about four weeks long. It starts with a decision, not a technology.
Week 1
Name the gate.
One place where a judgment gets made and a bad yes is expensive, and one precise question about it.
Week 2
Build the smallest thing that makes the question concrete.
An evidence record, a scoring sheet, a case that can be replayed. Anything built on invented data says so, everywhere it appears.
Week 3
Produce one number.
How far apart the reviewers landed. How much of the file was actually used. How often a structured record changed the call.
Week 4
Say what the number means, and what it does not.
Then decide: continue, narrow, redesign or stop. Stopping after a clear answer is a result, not a failure.
A product is what an investigation leaves behind when the number holds up.
Every result carries its rung.
Software that runs is not the same as a method that works. Everything we show is labeled with how far up this ladder it has climbed.
Illustrative demo
- It shows
- The workflow can be shown, on data labeled as invented.
- It does not show
- Accuracy, adoption or demand.
Checked prototype
- It shows
- The arithmetic, the edge cases and the reproducibility have been checked.
- It does not show
- That its assumptions describe your people or your cases.
Reviewed method
- It shows
- Someone qualified has examined the measures and the evaluation design.
- It does not show
- That it helps in practice.
Pilot
- It shows
- It has run on real cases, under stated conditions and safeguards.
- It does not show
- That it holds everywhere, or for good.
In use
- It shows
- A defined use, with review, monitoring and a plan for when it fails.
- It does not show
- That it is safe in a new setting or on a new model.
Figure 1 on the home page sits on the first rung, and says so.
Three things a good-looking result can hide.
Agreement is not accuracy.
Five reviewers can agree and all be wrong. A process that removes disagreement has not yet shown that it removed error.
Fast is not good.
A review that takes half the time may have skipped the part that mattered.
A convincing demo is not evidence.
It shows the software runs. It does not yet show that the method describes the world.
If it cannot end in a number, we do not do it.
Every piece of work here ends in a measured number about a judgment: how far apart two reviewers land on the same file, how often a structured record changes the call, what the variation costs. Work that ends in “your team uses more AI now” is somebody else’s work.