A draft came back with a scene in it that hadn’t occurred.
It was small. A reconciliation I supposedly ran against a calendar, a couple of sentences, entirely plausible. It was the obvious thing a competent operator would have done in that situation, which is precisely why it was there: the analysis had hit a gap in what I’d actually said, and it filled the gap with the most reasonable available answer.
It survived a full editorial pass. It was caught because I was the person who knew it hadn’t happened, and for no other reason.
That’s the whole argument I want to make, and I’d rather demonstrate it than assert it. I work with these tools every day. I’ve built four applications with them, and I’d take a serious pay cut before I’d give them up. They are also confidently wrong in specific, repeating, identifiable ways, and every one of those failures in my own work has been caught by knowing the business rather than by running more analysis.
Fabrication doesn't look like fabrication
The version everyone worries about is the outlandish claim, the invented statistic, the citation to a paper that doesn’t exist. Those are easy. They announce themselves.
What actually happens is subtler and much harder to see. The analysis reaches a point where it needs a fact it doesn’t have, and instead of stopping, it supplies the most plausible one. Plausible is the operative word. The invented detail is never bizarre. It’s the thing that should have happened, the step a sensible person would have taken, and it slots into the surrounding argument so neatly that it reads as more true than the material around it.
I now treat that neatness as the warning. If a detail fits too well, particularly one that supplies something I never provided, it goes on the list to verify. The tell isn’t implausibility. It’s convenience.
Four ways it went wrong in one working session
Over a few days of analytical work on a business I’d run for over a year, four failure types across three episodes — two of the four came out of the same piece of bad reasoning, which is itself worth noticing. I’m describing them by type rather than by figure, because the types are what transfer.
An invented specific. The one above. A concrete operational detail, supplied to fill a gap, indistinguishable from a real memory by anyone who wasn’t there.
A causal theory the source data contradicted. A confident explanation was built for why a particular change had caused a shift in the business: mechanism, reasoning, the lot. It was coherent and it was wrong. The underlying monthly data showed the shift hadn’t happened when the theory required it to happen. The theory had been built from summary figures and never tested against the detail underneath them.
An aggregate attributed to a single cause. A cost line had increased, and the entire increase was attributed to one thing. Most of it was that thing. A meaningful portion wasn’t. The resulting figure overstated a result by several points, in a direction that flattered me, which is the direction I’m least likely to check.
A conclusion from a metric nobody could reconstruct. A number in a reporting system was reasoned from at length. Nobody could reproduce its formula. Two separate conclusions came out of it and both were wrong, including the causal theory above, and they were only identified as wrong when the same question was rebuilt from raw data.
That overlap is the useful part. A single unverifiable input didn’t produce one bad answer, it produced a small family of them, all internally consistent with each other and none of them checkable against anything. Bad inputs don’t stay contained.
Every one of those was caught by domain knowledge. Not one was caught by more analysis, better prompting, or a second opinion from another model. They were caught because I knew what had actually happened in that business, and the analysis didn’t.
Why this is a judgment problem rather than a technology problem
These tools don’t know the difference between a number that measures something and a number that looks like it measures something.
I’ve written elsewhere about a consultation show-up rate that read 24% and was measuring administrative diligence rather than client behavior. The system recorded bookings automatically, and recorded completions only when someone remembered. Any analysis of that figure, however sophisticated, would have produced a confident account of a broken intake process. The data was internally consistent. The problem was upstream of the data, in what it was counting.
There is no amount of computation that finds that. You find it by knowing that the number moved when nothing changed, which requires knowing what changed.
That’s the first step of the method I use — the TAG method — See the Truth. Implement Actions. Achieve your Goals. Establishing what’s true is not the same as analyzing what you have. Analysis takes your figures as given and works from them. Establishing truth asks whether the figures describe what their labels claim, which is a question about the business rather than about the data, and it’s the question these tools cannot ask on your behalf.
Four checks
I use these on any analysis I didn’t produce myself, including my own team’s, including my own from six months ago.
Which specifics can I personally confirm? Go through the concrete details — events, sequences, actions attributed to named people — and mark each one you can verify from your own memory or a document. Anything unmarked is a candidate. Pay particular attention to details that make the argument work better than the facts you supplied would have.
Was the theory tested against the detail, or built from the summary? A causal explanation derived from aggregates will often survive contact with those aggregates and die immediately against the monthly data underneath them. Ask which one it was built from.
Does this attribute a whole effect to one cause? Almost nothing in a business has one cause. When an analysis assigns an entire movement to a single factor, it’s usually because that factor was the one under discussion, not because it did all the work.
Can I reconstruct every metric it relied on? If you can’t reproduce a number from source data, you can’t reason from it. Drop it. Two of my worst conclusions came from a figure nobody could rebuild.
What this actually means
The pessimistic reading is that you can’t trust the output. That’s not my experience. I get an enormous amount of value out of these tools and I’d be much slower without them.
The accurate reading is narrower. They collapse the cost of producing analysis to almost nothing, and they don’t collapse the cost of knowing whether the analysis is about the real business. That second thing is still expensive, still slow, and still made of time spent in the actual operation, watching what happens and remembering what you saw.
If you’ve run something for a year, you hold a set of facts nobody can retrieve and no model can infer. That was true before any of this and it’s more valuable now, because the analytical layer on top of it went from scarce to free, and the thing underneath it didn’t move at all.