An AI hallucination is output that reads as confident and fluent but is factually wrong, made up, or unsupported by any real source — and it happens because AI language models are built to predict the next plausible word, not to verify whether what they're saying is true.
That definition matters more than it sounds. A hallucination isn't a glitch or a bug that a patch will remove. It's the natural output of a system doing exactly what it was designed to do. When you ask a chatbot for a summary, a citation, or a statistic, it generates a sequence of words that fits the pattern of your request. Fluency is the goal. Accuracy is a side effect that usually happens, but not always. That gap between "sounds right" and "is right" is the whole problem.
Why the model can't check itself
To understand the cause, picture how these tools generate text. They work one token at a time — a token being a chunk of a word, roughly. At each step, the model looks at everything so far and calculates which next token is most probable. Then it picks one, appends it, and repeats. There is no separate step where the model pauses, looks up a fact, and confirms it before continuing.
That missing step is the core issue. A model has no built-in fact-checking layer, no internal encyclopedia it consults, and no way to mark a sentence as "unverified." It has patterns learned from training data and a statistical sense of what usually comes next.
When the pattern is strong — say, the opening lines of a famous speech — the output tends to be accurate. When the pattern is weak or the question is obscure, the model still produces something, because producing nothing isn't what it does. It fills the gap with the most plausible-sounding continuation.
Where the wrong answers come from
Three things feed hallucinations. First, training data gaps: if a topic was rare, outdated, or absent in the material the model learned from, it has no reliable pattern to draw on. Second, noise in that data: conflicting claims, satire, fiction, and errors all get absorbed as patterns, and the model can't always tell them apart. Third, the pressure to answer: when you ask a specific question, the model is nudged toward a specific-sounding reply, even if the honest answer is "I don't know."
Here's a concrete example. Ask a chatbot to name a study on remote work productivity and give the author and year. It may produce a tidy citation: a plausible researcher name, a real-sounding journal, a year that fits. Every element looks normal. But the study may not exist. The model didn't lie — it generated the shape of a citation, because that shape is what your question called for. This is why citation requests are one of the most hallucination-prone things you can ask an AI tool to do.
How to spot one
Fluency is not evidence. In fact, the more polished and confident a passage sounds, the more carefully you should check any factual claim inside it. A few practical habits help. Ask for sources and then verify that the sources exist — search the title, the author, the publication. Ask the same question a second time in a fresh conversation; if the model gives a different specific answer, that's a warning sign. And treat any number, date, quote, or legal or medical claim as unverified until you've confirmed it somewhere outside the tool.
It also helps to know what kind of task you're doing. Summarizing text you pasted in is relatively safe, because the material is right there. Brainstorming names or drafting a friendly email is low-risk. Recalling obscure facts, citing research, or giving precise figures is where hallucinations cluster.
Why you can't fully eliminate them
The honest limit is this: no amount of careful prompting removes hallucinations entirely, because the mechanism that produces them is the same mechanism that makes the tool useful. You can reduce them — by giving the model source text to work from, by asking it to say when it's unsure, by keeping questions narrow — but you can't switch them off. Tools that add web search or document grounding lower the rate, not to zero. For anything that carries real consequences, a human check is the last step, not an optional one.
This is also why the same tool can feel brilliant on one task and untrustworthy on the next. It isn't inconsistent in its design; it's consistent in a way that only sometimes lines up with the truth. Once you internalize that, you stop asking "is this AI reliable?" and start asking "is this the kind of question where reliability is likely?"