No True Understanding
Fluency in your vocabulary is not competence in your job.
This page covers one specific way AI gets things wrong at work, and what to do about it.
It runs in order. What goes wrong, why it happens, where you'd notice it on an ordinary day, who takes the blame, roughly what it costs, and the check that catches it. Then one thing to try this week.
The dollar figures are estimates, not measurements. The assumptions behind each one are printed right there, so you can swap in numbers that fit your job. Anything actually measured carries an OBSERVED tag.
What is no true understanding?
The model has learned the statistical relationships between words in your domain — which terms co-occur, what sentence shapes are typical, what tends to follow what. That is genuinely a great deal of structure, and it produces output that is correct far more often than chance.
But it is structure over language about the thing, not over the thing. There's no internal model of an obligation, a payment, a deadline, or a consequence. So the failure isn't random — it's concentrated exactly where language patterns and domain reality diverge. Edge cases. Exceptions. Situations where the usual phrasing implies the wrong outcome.
The model performs best on the cases you'd catch anyway, and worst on the ones you needed help with. That inverted difficulty curve is what makes this mode so costly: reliability on the easy 90% builds the trust that gets spent on the hard 10%.
What do people assume?
That a system using your domain's terminology correctly has some grasp of what the terminology refers to. That because it says "letter of credit" in the right places, it knows what a letter of credit does.
Correct usage and correct comprehension look the same from outside. This is the most expensive equivalence in the book, because it's the one that makes people stop checking.
Where does it show up at work?
A contracts analyst asks an assistant to summarize the termination provisions of a vendor agreement. The summary is accurate, well-organized, and reads like it was written by someone who does this for a living.
It omits that the termination-for-convenience clause is subordinated to a minimum-purchase commitment elsewhere in the document. The two clauses are eleven pages apart and don't share vocabulary. Nothing in the language of either signals the interaction.
The company gives notice. It still owes the balance of the commitment — in a mid-six-figure agreement, somewhere in the $200K–$400K range.
(Composite scenario. Constructed from the structure of a common commercial failure, not a reported incident. Range shown rather than a point estimate because a point estimate here would be false precision — the exact figure depends entirely on contract size and remaining term.)
Who carries the downside?
Vendor: none. Executive: approves the notice based on the summary. Manager: signed off. You: produced the summary. In the post-mortem, the question is not "what did the tool miss." It is "who read the contract."
What does it cost?
[MODELED — not reported]
ASSUMPTIONS Consequential docs summarized: 30 / year Rate with a material interaction the summary misses: 8% (~2.4 / year) Rate caught downstream anyway: 60% Incidents reaching a decision: ~1 / year Cost per incident: $40,000 – $200,000 (value of the missed obligation or forgone right)
Annualized exposure: ~$40,000 – $200,000
How do you control for it?
Read the source document for anything consequential. Use the summary as an index, never as a substitute. Specifically: identify every cross-reference, every defined term, and every clause that conditions another clause — the places where meaning lives between sections rather than inside one.
CONTROL COST Consequential documents: 30 / year Additional human read: 90 minutes each Annual: 45 hours Fully loaded rate: $85 / hour (senior)
Annualized control cost: $3,825
What should you do this week?
RECOMMENDATION
Write down the five places in your work where meaning depends on two things being true at once — a clause that modifies another clause, a rate that applies only under a condition, an approval that expires. These are your interaction points, and they are where this failure mode lands every time.
Then state the rule out loud in your next team meeting: summaries route you to the document; they don't replace it. You have just defined a control, in public, before anything went wrong. That is a materially different position than defining it afterward.
Evidence
RESEARCH Language models learn distributional structure over form without grounding in referents.
ANALYSIS The inverted difficulty curve — strong on routine cases, weak on the exceptions where help was needed — is the author's framing.
ANALYSIS The $340,000 contract vignette is illustrative and composite. Must be labeled as such in the final text. Do not present as a reported incident.
This is one of 24 failure modes. The book gives you all of them — plus the controls that catch each one and a 90-day plan to prove you ran them.
Preorder the Book