Method

Judge the signal by the decision it feeds.

The homepage asks two questions: is the model right, and is it being used right. This is how I answer them, in roughly the order the work uses each idea.


Start with the decision, not the model

A signal is only worth what it does for the decision it feeds. So the first questions aren't about the model at all. What decision does this output inform? What does a miss cost, and what does a false alarm cost? How often does the thing you're looking for actually happen?

That last one, the base rate, does more work than it gets credit for. A system flagging a rare event can be wrong most of the times it fires, even when it's very good, simply because the event is rare. Whether that's acceptable depends on the two costs, not on the accuracy figure.

Weather forecasting runs on this idea. A forecast has no value on its own. It has value when someone decides differently because of it: whether to fly, whether to protect, whether to wait. The same is true of a credit score, a fraud flag, or an agent's draft.

Accuracy is a weighted average

A vendor, or your own team, isn't wrong to call a system 90% accurate. But that number is a weighted average: close to 100% on one slice of cases, close to a coin flip on another, blended in whatever mix showed up in the data used to test it.

Which slice is which isn't something you can tell by looking. It's a property of what the model was trained on, not of what looks routine to your business. The mix you actually see in production can be a completely different pie than the one it was tested on, same slices, different proportions. If the weak slice is bigger in your world, the accuracy that matters is nowhere near 90%, and nobody lied to get there.

What got tested
Blends to 90% reported
What production looks like
Blends to 70% real

This is the "is the model right?" question. It's complicated rather than complex: hard, but it yields to method. You measure it on your own history, slice by slice, rather than taking the blended number on trust.

Human judgment is a dial, not a switch

Once you know where the system is reliable and where it's closer to guessing, the design question is how much human judgment each case gets. Not a single cutoff where the machine takes over, but a dial that turns with the system's confidence, the cost of being wrong, and the stakes of the particular case.

How much human judgment a case gets turns with the system's confidence Cases spread from low confidence to high. The shading shifts gradually from more human judgment, where the model is closer to guessing, to less, where it is reliably sure. There is no single cutoff. less sure more sure more human judgment the model is closer to guessing less human judgment the model is reliably sure
Every case a system handles sits somewhere on this curve. A single accuracy number averages across all of it.

Setting that dial well is how scarce expertise gets spent where it's earned, and how routine work stops needing the same oversight as the hard cases. It's also where "is it being used right?" gets answered, and that one is complex rather than complicated. The answer depends on the people, the process, and what happens after deployment, so it doesn't yield to method alone.

The answer has to arrive before the decision

Timing is a design property, not an afterthought. A score that posts overnight, or a flag that lands in a queue someone works the next morning, can be perfectly accurate and still change nothing. If the signal arrives after the decision has gone out, it's a record, not a control. So part of the work is mapping when each decision actually gets made, and checking that the answer lands in time to change it.

Signals feed signals

AI output rarely stops at a person. It becomes the input to the next system, the next model, the next decision, and those decisions become inputs too. Errors don't stay where they start. A modest miss rate upstream can compound two steps later, in a place nobody thought to look. Judging a system in isolation misses that, so the evaluation follows the signal to the decisions it actually feeds, not just the first one.

Evidence has a shelf life

Evidence is about a specific system: this model, this prompt, this data, this use. Change any of them and the evidence describes something that no longer exists. Systems drift, get retrained, and get pointed at problems they were never tested on. Staying decision-grade means knowing which changes invalidate what you know, and checking again before the numbers quietly stop meaning what they meant.

Mechanisms, not the destination

Evaluations, benchmarks, a human in the loop, whatever the mechanism, are how you keep pace as the regulations, the models, and the data underneath them change. They're mechanisms for keeping up. The destination is a better decision.

Let's talk

A 30-minute call, no pitch.

info@grayduckpartners.com