You Turn the Page, and Forget What You Know
The bias that makes casual use of AI more dangerous than you think
Some of this article is also covered in a video on Exploring ChatGPT’s YouTube channel - take a look and subscribe!
In April 2002, Michael Crichton — novelist, screenwriter, Harvard-trained physician, creator of Jurassic Park — gave a talk called “Why Speculate?” and buried in it a small observation that has aged remarkably well. He named it after the Nobel Prize-winning physicist Murray Gell-Mann, with deliberate self-irony: he had once discussed the phenomenon with Gell-Mann, and “by dropping a famous name,” he said, he was implying greater importance than the idea deserved.
The effect itself is disarmingly simple.
You open a newspaper to an article about something you know well, like your industry, your technical specialism or your profession. You read it and find it riddled with errors. Causes and effects reversed. Nuance flattened into nonsense. Crichton called these “wet streets cause rain” stories, and noted that serious publications are full of them. You put the paper down, mildly exasperated. Then you turn the page to an article about foreign policy, or macroeconomics, or public health (topics where you might have less depth) and you read it with the same credulous attention you’d give a trusted colleague. You forget, entirely, what you just witnessed in the previous article.
Crichton framed this as a media problem. It isn’t, or at least it isn’t only that. It’s a cognitive pattern, one that travels with us wherever we encounter authoritative-sounding output in territory we can’t fully evaluate. And it has found, in AI, the most capable host it has ever had.
Why the Confident Voice Switches Off the Critical Mind
The Gell-Mann effect rests on two things working together:
genuine expertise is domain-specific, so the scepticism we've built through years in our own field doesn't port across to the next one; and
credibility earned in one area bleeds implicitly into adjacent ones, through the same cognitive shortcuts that make everyday life navigable.
Crichton pointed out the troubling asymmetry here by invoking the legal principle falsus in uno, falsus in omnibus, or untruthful in one part, untruthful in all. Courts apply this logic routinely, but readers don’t. In ordinary life, he noted, if someone consistently exaggerates or misleads you, you discount them across the board. But with media, and now with AI, we extend a strange special exemption. We see the errors where we can spot them, and then proceed as if they’re localised anomalies rather than evidence about the underlying quality of everything else.
The Art of Asking Questions is a reader-supported publication. To support my work, please consider becoming a paid subscriber.
AI Is the Best Gell-Mann Machine Ever Built
LLMs produce exactly the surface features - fluency, structure, measured tone - that trigger trust transfer, and they do it uniformly across every domain, whether or not the underlying content is sound.
Consultants encounter this directly. In areas of genuine expertise, the errors are often visible immediately. The scepticism fires, the output gets interrogated, and the error gets caught. That critical friction is a feature, and it’s what working well with AI actually looks like.
The problem is what happens outside the zone where we are sufficiently knowledgeable to spot errors. A developer who spots AI-generated code errors instantly may consult the same model on, say, employment law implications, or the epidemiology behind a health claim in a client report, and find the output entirely convincing. The scepticism that fires reliably in one domain simply doesn’t port across.
The real danger is the selective nature of where our scepticism fires. We catch what we can catch, and trust the rest almost implicitly, because the output doesn’t look any different.
The Antidote is Portable Scepticism
The antidote to Gell-Mann Amnesia cannot be blanket distrust. That would make everyday life impractical and tools like generative AI unusable. The goal is to make your scepticism portable.
Start by using your own error-catching experience as a calibration baseline. If you notice flawed output in areas you can verify, treat that error rate as your prior for everything else the same source produces. The quality of what you can’t evaluate isn’t magically higher than the quality of what you can.
Then, learn to identify the moment you stop being able to evaluate the output. That threshold, where you’re reading with understanding rather than with expertise, is where your scrutiny should increase. Make it explicit: what, specifically, would you need to know to properly assess this section? If the answer is “quite a lot that I don’t currently have,” that’s the signal to slow down.
Crichton concluded in 2002 that “the only possible explanation for our behaviour is amnesia.” He was talking about newspapers, but with AI there is no byline to track and no visible seam between the domains it handles well and the domains where it reverses cause and effect. Those best placed to use these tools are the ones who have mapped, honestly, where their own knowledge runs out.
Knowing the effect exists is the easy part. Knowing exactly where it applies to you, in your practice, on the work you actually do, is what matters.
The AI Output Validation Funnel
The arguments above lead to a practical question: if your scepticism is strongest where you need it least, what do you actually do about it? Here’s the pipeline that Andrea uses in his day job as a consultant.
The core principle is simple:
Your QA effort should increase in proportion to how far the content sits from your own expertise.
The steps below are layered so that each one catches a different type of error, and no single step covers everything - which is the whole point. We’ll start with a visual summary, and the full explanation follows.
You can also watch a video covering the AI output validation funnel. Check out Exploring ChatGPT’s YouTube channel and subscribe!
1️⃣ Flag your expertise boundary. Before anything else, go through the AI output and assess which sections you can evaluate from your own knowledge and which you can’t. Many people skip it. The sections you can’t properly assess are where the rest of the pipeline matters most. Be honest with yourself here: understanding a passage is not the same as being able to evaluate whether it’s correct.
2️⃣ Cross-check and challenge with a second model. There is an additional complication worth highlighting. AI outputs are often mostly right, with errors embedded in surrounding claims. This is in some ways worse than wholesale fabrication, because the correct material lends credibility to the incorrect material in the same passage.
So I recommend you take your AI output and run it through a different model from the one that generated it. You are not looking for the “right answer” from a different AI tool, but for signals of divergence. Where two models confidently say different things, you have found a claim that surely needs in-depth verification. Where they agree, you have slightly higher confidence but not certainty, because they often share a lot of training data and can reproduce each other’s blind spots. The value here is in surfacing disagreements you wouldn’t have spotted yourself.
To get the most out of this, don't just ask the second model the same question, but prompt it to actively critique specific features of the original output:
“Here’s a passage about [topic]. What’s wrong with it? What claims are unsupported, overstated, or likely to be challenged by someone with deep expertise in this area?”
[Paste the generated text]
This works surprisingly well for the “mostly right with embedded errors” problem described earlier. The AI won’t catch everything, but it will often flag the structural weaknesses: reversed causation, overstatement of effect sizes, missing caveats, conclusions that don’t follow from the evidence presented.
Here's an example where I ran a Claude-generated passage through ChatGPT and it flagged reversed causation and overstated certainty:
So the short verdict is: the passage is readable and mostly accurate at a high level, but it would be challenged for being too linear, too certain, and occasionally wrong on specifics.
https://chatgpt.com/share/69ac53db-5448-800f-ae13-cc921676b6e6
You can also feed the critique back into the original model and ask it to respond: advanced reasoning models working in dialogue can surface issues that neither catches alone.
3️⃣ Verify sources for specific claims. Any named statistic, attributed quote, cited study or specific factual claim (dates, figures, institutional positions) needs to be checked manually against a primary source. This is the step most people skip when the output “feels right”. AI models fabricate citations with complete fluency, so you must use web search, Google Scholar, or similar sources for this step, not another AI model.
4️⃣ Audit logic and structure. Read the output specifically for argumentative structure:
Are causes and effects in the right order?
Does the conclusion actually follow from the evidence presented?
Are there implicit assumptions doing heavy lifting?
This is the “wet streets cause rain” check from Crichton’s original talk, and you can do it yourself for most content regardless of domain expertise, because you’re evaluating the reasoning rather than the subject matter.
5️⃣ Ask an expert to spot-check high-stakes work. If the content is going to a client, into a publication or anywhere with reputational consequences, get a domain expert to review the sections outside your expertise. This doesn’t have to be a full review: even a 15-minute read of the flagged sections by someone who works in that area will catch things the entire AI pipeline won’t.
This all takes time. The honest answer is that quality assurance is more important than ever when using generative AI. Andrea and Cécile speak about that in detail here:
But, clearly, efforts should scale with the stakes:
A quick internal note might only need the first two or three steps.
A client deliverable or published article needs the full pipeline.
That understanding keeps things practical rather than aspirational, and it mirrors how most professionals already calibrate their review effort for human-produced work.
The difference is that AI output needs this calibration made explicit, because the surface quality doesn’t vary between the sections it got right and the sections it got wrong.
👇 Steal My Toolkit
The newsletter covers ideas, frameworks and honest takes on consulting practice. The paid tier is where those ideas become usable: you get templates you can open in a project, prompts you can run today, and research tools built from over a decade of fieldwork.
Paid members get instant access to:
Consulting with AI Series: the custom AI prompts I use in my own management consulting work, explained and ready to adapt.
Full Course: Asking Questions Like a Pro - A Practitioner’s Guide to Mixed Methods Research ($199 value): practical research techniques refined across 10+ years of consulting engagements, covering design, facilitation and analysis.
Practical Guides, Handbooks, Workbooks and Templates: every premium resource from paid posts.
Cozora discount (up to $360 off): up to 50% off a Cozora subscription to learn AI from practitioners who use it in real work.
Free subscribers aren’t left out either.
Leader Tools Cards: 10% off with code ANDREA10 at leadertools.co/ANDREA10
Cozora welcome discount: 10% off with code WELCOME10 at cozora.org
Demystifying Academic Research - A practical guide for navigating academic research without a PhD: Available for free (pay what you can, $0+) via Gumroad
Disclosure: I earn a referral fee if you purchase through my links, at no extra cost to you.






Nice post. The idea of “portable skepticism” is useful, especially the point that our critical thinking doesn’t transfer cleanly across domains.
I like how situational this is in practice. Most of the time people aren’t trying to validate everything, they’re deciding where to spend attention. The real move is recognizing when something shifts from interesting to consequential, because that’s when the level of scrutiny changes.