EVIDENCE / HUMAN JUDGMENT

How to fact-check AI answers—even when several chatbots agree

The Illuminated Shore editorial team · Original companion guide · October 2, 2026

Check important claims against original evidence. Several chatbots repeating an answer may share a source or an error. A useful second check uses a different route to the truth: a document, a calculation, or a test.

Confidence is a style; support is a relationship

An AI answer says a classroom tool “improved learning by 20 percent.” The sentence sounds precise. Before using it, you need to know what improved, compared with what, in which students, and over what period. A link to a real study is only the beginning.

Separate a factual claim from an interpretation or recommendation. “This trial measured a higher test score” and “every school should buy this tool” require different kinds of support. You can agree with the first while rejecting the second.

Follow the claim through four checks

  1. Extract the exact claim. Write the number, population, outcome and time period as separate items. Do not let a fluent paragraph blur them together.
  2. Open the original source. Prefer the study, official record or original announcement. A summary may omit conditions or describe an earlier version.
  3. Match the evidence. Find the passage, table or data supporting the claim. Check units, publication dates and whether the result is measured, modeled or surveyed.
  4. Test the conclusion. Ask what the evidence leaves unresolved. Remove or narrow a sentence that reaches beyond it.

For a current product feature, the current provider documentation is usually more relevant than a remembered answer. For a calculation, recompute it. For software behavior, try a representative case. Choose the check that can expose this particular mistake.

A small number with a large difference

CONSTRUCTED EXAMPLE · NOT A RESEARCH RESULT

40% becomes 50%. What changed?

If 40 of 100 people finish a task in one group and 50 of 100 finish in another, the difference is 10 percentage points. Relative to 40%, it is a 25% increase: (50 − 40) ÷ 40.

Neither figure establishes that a tool caused the change. That depends on how the groups were formed and what else differed.

Ask an assistant to explain both calculations and identify the missing design information. Then check the arithmetic yourself. The goal is an inspectable answer, not a more emphatic one.

Why another chatbot is not automatically another source

Two assistants may rely on the same article. Three reviewers may read the first reviewer’s mistaken summary before forming their own view. Different role names do not establish independent evidence.

Research on sycophancy found that the tested assistants sometimes favored responses matching a user’s expressed beliefs. The study is useful evidence of a failure mode; it is provider-authored research from 2023, not a current accuracy rating for all products. [1]

For a consequential claim, let a checker work from the original source before showing the proposed conclusion. Ask for a reason to accept or reject each claim. Do not require disagreement: an invented objection is also an error.

Keep a compact claim ledger

A reusable review instruction

For each important factual claim, list the original source, the supporting passage or table, the source date, and what the source does not establish. Mark unsupported claims as unresolved. Keep proposals separate from findings. Do not invent a citation to complete the table.

This instruction improves the review’s structure; it does not guarantee the assistant follows it. Open the sources yourself when the stakes justify it. If a decision could materially affect someone’s health, rights or finances, bring the evidence to the relevant qualified professional.

Both books ask readers to notice the gap between an attractive signal and a useful result. The newer book applies that question to AI reviewers and teams; the original follows it through attention, screens and evolution.

Sources & limits

  1. Sharma and colleagues, Towards Understanding Sycophancy in Language Models (2023)

    Provider research testing particular assistants and tasks. Agreement with a user is not independent factual verification.

Research links checked October 2, 2026. These sources do not endorse the books or this site. Exercises and practical suggestions are original companion material, not research findings or manuscript excerpts.