Uncertainty Miscalibration
A stated 90% confidence is a linguistic performance, not a probability.
This page covers one specific way AI gets things wrong at work, and what to do about it.
It runs in order. What goes wrong, why it happens, where you'd notice it on an ordinary day, who takes the blame, roughly what it costs, and the check that catches it. Then one thing to try this week.
The dollar figures are estimates, not measurements. The assumptions behind each one are printed right there, so you can swap in numbers that fit your job. Anything actually measured carries an OBSERVED tag.
What is uncertainty miscalibration?
Distinct from Mode 4, which concerns tone. This concerns explicit numerical self-assessment.
A model asked to state a confidence level produces a plausible-sounding number. There is no operation connecting that token to the model's internal distribution over answers. Verbalized confidence and actual accuracy can be substantially decoupled, and the decoupling tends to run in the direction of overstatement.
Worse for operational purposes: the numbers cluster. Ask for confidence across many claims and you get a narrow band — lots of 85s and 90s — regardless of how much the true reliability varies. A confidence score with no variance carries no information, and a score that doesn't discriminate is worse than no score, because it gets used for triage.
What do people assume?
That when a model attaches a number to its confidence — "I'm about 85% sure" — the number means something in the frequentist sense. That across many such claims, roughly 85% will be correct.
The number is generated text. It is not read from an internal probability.
Where does it show up at work?
A risk analyst builds a review queue that routes anything below 80% model confidence to human review. It looks like a well-designed control with a sensible threshold.
Nearly everything scores above 80. The queue is empty. The control reports green and reviews nothing.
Who carries the downside?
Vendor: none. Executive: believes there is a functioning control. Manager: reports on it. You: designed the threshold. This is worse than having no control, because the organization has now stopped looking.
What does it cost?
[MODELED — not reported]
ASSUMPTIONS Items routed by confidence score: 10,000 / year True error rate in "high confidence" bucket: 2% (200 errors) Errors that would have been caught by working triage: 60% (120) Cost per uncaught error: $500 – $5,000
Annualized exposure: ~$60,000 – $600,000
The largest modeled exposure in Part II — because a broken control operating at scale is more expensive than no control at all. No control means people stay alert. A green dashboard means they don't.
How do you control for it?
Validate the score before trusting it. Sample outputs across confidence bands and measure actual accuracy per band. If accuracy doesn't separate across bands, the score is not usable for routing and the threshold must be replaced with a rule based on observable features — document type, dollar value, customer tier, regulatory flag.
CONTROL COST Quarterly validation: 4 / year Time per validation: 12 hours Annual: 48 hours Fully loaded rate: $85 / hour
Annualized control cost: $4,080
What should you do this week?
RECOMMENDATION
Find every place your organization routes work by a model-reported confidence score. Ask one question: has anyone measured whether accuracy actually differs across the bands?
If nobody has, you have found a control that is reporting green while doing nothing — and you found it before the auditor did. That is the highest-value finding an individual contributor can produce in a governance review, and it requires no permission to look for.
Evidence
RESEARCH Decoupling of verbalized confidence from actual accuracy.
ANALYSIS Confidence-score clustering and the resulting loss of discriminative power.
ANALYSIS The "broken control is worse than no control" argument is the author's. Retain the label — it is the strongest claim in the chapter and must not be dressed as research.
This is one of 24 failure modes. The book gives you all of them — plus the controls that catch each one and a 90-day plan to prove you ran them.
Preorder the Book