You cannot tell whether an AI answer is a hallucination just by reading it — confident, fluent, well-formatted text is exactly what a language model produces whether it is right or wrong, so the only reliable test is to check the specific claim against a primary source outside the AI.
An AI hallucination is a statement a model generates that sounds plausible and is stated with confidence but is factually wrong or invented, and because the model has no internal "true/false" flag it can expose, the text itself gives you no signal.
That means your job is not to detect the lie by reading closely. It is to decide, before you trust the answer, whether the claim is the kind of thing you can go and verify.
The mechanism matters here, because it explains why confidence is worthless as a clue. A language model generates text by predicting what word or phrase most plausibly comes next, based on patterns it absorbed during training. It is optimising for fluency and plausibility, not for correspondence with reality.
When it does not have a solid pattern to draw on — an unfamiliar citation, a precise date, a niche product's plan names — it still produces something that fits the sentence's shape. That is why hallucinations cluster around specifics: names, numbers, dates, quotes, and citations.
A model is far more likely to invent a plausible-sounding study title than to get the general idea of photosynthesis wrong. So the useful question is never "does this sound right?" but "is this claim the type that has a single checkable answer?"
Here is a concrete worked example. Suppose you ask an AI assistant for the pricing of a project-management tool so you can budget for a small team. The answer comes back with a clean table: a free tier, a per-user monthly plan, and a business tier, all with specific dollar figures and a confident note that the business plan includes advanced automation.
The prose is tidy and the numbers look reasonable. Now apply the test. Is this claim checkable?
Yes — pricing has one authoritative source, the vendor's own pricing page. Is the AI's version likely to be stale or invented? Very likely, because pricing changes often and the model's knowledge has a cutoff.
So you open the vendor's page and compare. If the plan names and figures match, you have a verified answer. If they do not, you have caught a hallucination without needing any special skill.
Note what you did not do: you did not reread the AI's paragraph looking for tells. There were none to find.
For a second example with a different failure mode, imagine asking the same assistant to explain why some teams find an AI coding tool brilliant while others find it useless. That is a reasoning question, not a lookup. There is no single page that settles it.
Here the checkable-claim test tells you something different: the answer is not verifiable in the same way, so you should treat it as a hypothesis to think with rather than a fact to cite. You can still sanity-check the logic — does the explanation account for the cases you have seen?
— but you should not repeat it as established fact. Our internal AI tool database, which records pricing and capability snapshots for 360 AI tools at verification time, exists precisely because the checkable claims are the ones that go stale fastest. That is a reason to check primary sources, not a substitute for doing so.
Now the limits, stated honestly. This approach fails in three situations. First, when the claim is genuinely unverifiable — predictions, opinions, or open research questions have no primary source to check against, so you are left judging the reasoning, which is much harder.
Second, when verification is expensive: checking a claim might take longer than the task was worth, and in that case the honest move is to label the claim as unverified rather than pretend you checked. Third, when the AI is right about something counterintuitive and your quick check is shallow — a single glance at a search result can mislead you as easily as the model did.
The rule that survives all three: before you rely on any specific name, number, date, or quote from an AI, identify the one source that would settle it, and go there. If no such source exists, downgrade the claim in your own notes.