# Confidence percentages are theater

**Summary:** A confidence score next to an AI answer looks like rigour and usually measures nothing. Here is what to ask your vendor instead.

**In short:** A number next to an answer feels like evidence, and almost never is.

**Published:** 2026-09-03

**Last updated:** 2026-09-15

Somewhere in your AI tooling there is a number. An answer appears on screen, and beside it sits 87%, or 0.91, or a small green bar labelled high confidence. It looks like evidence. It is usually the least evidential thing on the page.

This is not a complaint about dishonest vendors. The number is there because everyone expects it to be there. But a figure that nobody measured against anything is decoration, and in a regulated business decoration is expensive, because sooner or later somebody asks what it meant.

## Where the habit came from

Product teams have been scoring their own certainty for years. Jason Cohen makes the case in [Lost confidence](https://longform.asmartbear.com/confidence/ "A Smart Bear: Lost confidence"), from June 2026: the confidence column in a RICE-style prioritization framework asks you to estimate a probability you have already admitted you cannot know.

His fix is the part worth stealing. Replace confidence with asymmetry, and estimate two things you actually can: the worst case in time and money, and the best case in customer value. Both have checkable answers, where a probability of being right has none.

The argument is about roadmaps, but it transfers, and it transfers to a much bigger audience. A product manager fudging a confidence score affects a roadmap. A support system shipping an answer under a reassuring green bar affects a customer who believed it.

## Why it gets worse when a model is involved

A person at least knows they guessed. When the number comes out of a system, that knowledge disappears, and the figure arrives looking like an instrument reading.

Ask the question that settles it: what was this number compared against? For a real measurement there is an answer. The system was run against a set of questions whose correct answers were known in advance, it got some right and some wrong, and the rate is the number. For most confidence displays there is no such set. Nothing was scored, so nothing was measured.

That is the whole distinction, and it is worth being blunt about it. A confidence score is only as good as the thing it was checked against. No control set, no meaning.

## The strongest objection, which is a good one

Calibration is real, and dismissing it wholesale would be its own kind of theater.

A system that says 90% and is right nine times out of ten across a few thousand scored questions is telling you something genuinely useful, and the work to get there is serious engineering. The objection is fair: do not confuse a measured number with an invented one just because they render identically.

So the argument is narrower than "confidence scores are bad". The two look identical on screen and mean opposite things, and almost nobody asks which one they are looking at. Keep the number. Make the question routine.

## What we do instead, and what we would ask a vendor

Unless takes the other route, and it is stated on our own [trust page](https://unless.com/en/trust/ "Unless: trust and compliance") in plain terms: no black box, every output points back to its source, and if we cannot show our work we do not ship the answer.

In practice that means two things you can inspect rather than trust. Every decision leaves a per-decision audit trail carrying a timestamp, the decision path, the model used and the source cited. And accuracy is measured the boring way, against control questions with known control answers, in the [Test phase](https://unless.com/en/engine/test/ "Unless Engine: Test") of the loop, so the number that exists is one somebody scored.

We are not claiming this is the only defensible design. We are claiming it is the one we can show you, which is a different and smaller claim than a percentage implies.

If you are evaluating anything that puts a confidence figure on screen, four questions separate a measurement from a mood. What set of questions was this scored against, and can I see it? How often is it rescored? What happens to the number when the underlying content changes? And when the system is wrong, can I read back why it answered the way it did?

A vendor with a real number will enjoy those questions. That reaction is itself the answer.

## What you lose, and what you get

Giving up invented precision feels like giving up certainty. What you are really trading is a number that reassures for a record that survives a conversation with your auditor.

Deciding under uncertainty is the normal condition of running a regulated business. Nobody in your compliance function expects an interface to have removed uncertainty. What they expect is that when something goes wrong, you can explain what the system knew, where it got it, and who could have stopped it. A percentage cannot do any of that. A trail can.
