Why your AI makes things up so confidently
It gave you a price, a date and a source. It all looked right. You nearly published it. Checking by accident, you discover none of it exists: not the figure, not the study, not the link. And it stated all of it with the same confidence as everything else.
An AI makes things up because the way it is evaluated rewards guessing and punishes admitting ignorance. Like a student facing a hard question, it attempts a plausible answer rather than leaving a blank. The confidence is not a sign of certainty, it is the default style.
It is neither a lie nor a fault
It gave you a date. A reference. A function name that does not exist. All of it with perfect assurance, without a flicker of hesitation.
The first instinct is to wonder whether the tool is broken, or whether it is “lying”. Neither. Lying requires knowing the truth and choosing otherwise. Here, the machine does not know that it does not know.
And the underlying reason is more interesting than a simple technical fault.
The student who is better off guessing
The clearest explanation comes from the researchers who studied the question, and it fits into a comparison everybody has lived through.
Picture a marked exam: a right answer scores one point, a wrong answer zero, and a blank also zero. You reach a question whose answer you do not know. What do you do?
You have a go. Obviously. Leaving it blank guarantees zero, while guessing gives you a chance. A student who guesses systematically scores better than an honest one who leaves blanks.
Language models are graded in exactly that way. The vast majority of the evaluations used to compare them mark each answer right or wrong, with no credit for well-expressed uncertainty. The system therefore rewards guessing and punishes admitting ignorance — and models become excellent exam candidates.
Invention is not a bug that appeared by accident. It is the behaviour the marking encourages.
And why the confidence on top
One question remains: why say it with such assurance rather than hedging?
Because tone is not wired to certainty. A well-founded answer and a guessed one come out with exactly the same crisp phrasing. Nothing in the delivery tells them apart.
That is what makes the phenomenon so expensive for a beginner. On a subject you know, you catch the error in a second. On a subject you do not know — and that is precisely why you asked — you have no signal at all.
Spotting an invention in three seconds
You do not need to be an expert to catch the important ones. Three reflexes are enough, and they cost nothing.
Suspicious precision. The more precise a claim, the more it deserves a look. “Many sites have this problem” commits to nothing. “43% of sites have had this problem since March 2024” is exactly the kind of sentence that manufactures itself.
The link nobody clicks. A page address quoted in an answer is checked by clicking it. It is the fastest test in existence, and the one nobody runs. A page that does not exist is an invention caught in two seconds.
The follow-up question. Ask “what makes you say that, and where did you read it?”. A grounded answer gives you a source you can open. A guessed one suddenly gets much vaguer, or produces a source that also fails the click.
What does not work
Asking whether it is sure. It will say it is sure, with the same assurance. The question has no verifying power whatsoever.
Assuming a newer model settles it. They invent less often, and the researchers publishing them write themselves that the problem is not solved. Less often is not never.
Telling yourself you will notice. In your own field, yes. Everywhere else, no — and everywhere else is where you ask questions.
The right stance
An AI is not an encyclopaedia that is occasionally wrong. It is a tool producing plausibility, and plausibility is true most of the time.
So the right posture is neither permanent suspicion, which makes the tool useless, nor blind trust, which ends up publishing a wrong date on your site. It is one simple rule: anything precise and checkable gets checked before it is published. The rest — structure, phrasing, ideas — carries no such risk.
It is the same principle as the rest of the work: you do not change the tool, you set the frame around it.
> how many sites have this problem?