AI systems for decisions that matter

Better decisions and outcomes from your AI systems.

Every AI system or model sends your business a constant stream of signals, at a speed and scale no team can check by hand. Deciding what to do with them is where I work. That takes two things, done together: evaluation that judges the signal by the decision it feeds, and system design that acts on what the evaluation finds. The goal is a better decision, not a better evaluation.


Two questions. One complicated, one complex.

Any system making decisions for you raises both, and they fail in different ways.

1. Is the model right?

The one everyone asks first, and the one a vendor's accuracy number is usually answering.

  • "It seems to work" is the whole case. No defensible answer to how good is good enough, and nothing built in to watch it, so there's no way to tell whether it's still right.
  • The evidence came from somewhere else. A result demonstrated on the vendor's data, or on a different version of the system, is a claim about that system and not about yours.

2. Is it being used right?

Complex rather than complicated, which is why it gets overlooked. The answer depends on context and on what happens after deployment, so it doesn't yield to method alone. Almost nobody measures it, and it's where most of the real risk lives. An accuracy figure is an average across a population. You don't experience averages. You experience particular customers, particular files, particular exposures.

  • Every case gets treated the same. The system behaves identically when it's confident and when it's guessing, and the threshold deciding what it handles alone was inherited rather than chosen. Nobody can say what a miss costs compared to a false alarm.
  • The answer arrives after the decision. The score posts overnight, or the flag lands in a queue somebody works the next morning. A signal that lands after the decision has gone to the customer is a record, not a control.
  • It drifted, or it got pointed somewhere new. A system bought for one purpose gets aimed at another, and the numbers quietly stop meaning what they meant.

My work is for anyone who wants better outcomes from an AI or agentic system, and better takes different shapes. Sometimes it means risk that's documented and defensible to a regulator. Sometimes it means the business finally trusts the system enough to adopt it, before a costly build gets abandoned. Sometimes it means a team finally sees efficiency gains, because scarce human judgment gets spent where it's actually needed.

Read the details here →

What decision-grade looks like

Every failure mode above has an answer. When the answers are in place, four things are true.

How it goes

We agree what is at stake
What it decides, what being wrong costs, and which parts deserve real scrutiny, worked out in sessions with the people who built it and the people who rely on it.
I evaluate, then design around it
What the evaluation finds, the evidence behind it, and the architecture built to act on both. Understanding how a decision actually gets made takes ongoing conversations with your team, not a document review from a distance.
Your team can act on it
A plan to actually move the business outcome you came for, what changes first, in what order, and why, walked through live so the reasoning transfers instead of sitting in a document.

You work with me the whole way, with no handoff to a junior team.

What this is not

  • Not an "AI readiness" questionnaire or a strategy slide deck.
  • Not a vendor or model pitch.
  • Not a generic audit. It ends in decisions your team can act on, not observations they have to translate first.
  • Not a one-time scorecard of evaluations and benchmarks. Those are mechanisms for keeping up as the regulations, the models, and the data change, not the destination.

More about my background →

Take the next step

A 30-minute call to see if this fits. No charge, no pitch. If it is a fit, we take a longer session to work out what a good outcome is actually worth to you. The scope and the fixed price follow from that.

info@grayduckpartners.com