Almost everyone who works with AI eventually has this moment: the answer is written so smoothly that checking it never crosses your mind, and then you realize a piece of information in it was completely made up. This is called AI hallucination. In this article we explain why hallucinations happen, which kinds of work they are most dangerous in, and how a business can make fact-checking systematic.
Why do hallucinations happen?
Language models do not work like a database of facts. Based on patterns learned from the huge body of text they were trained on, they generate the most likely words to continue a sentence. Most of the time that likely continuation matches the correct information. But when the model lacks enough knowledge about a topic, it tends to produce the most plausible-looking answer rather than stay silent.
Typical situations that raise the risk of hallucination:
- Very specific or little-known topics (a small company's product code, a local regulation clause)
- Questions that need current information (prices, regulations or exchange rates that changed after the model's training date)
- Questions asking for exact figures, dates, sources and quotations
- Long, multi-step calculations
- A false assumption built into the question ("according to Article 12 of such-and-such law…")
Not every error is equally dangerous
Checking every task with the same rigor is both unnecessary and unsustainable. It is smarter to match the level of checking to the risk of the task.
| Risk level | Example task | Check |
|---|---|---|
| Low | Headline ideas, internal message draft, brainstorming list | A quick read is enough |
| Medium | Blog post, product description, customer email draft | Check figures, names and claims one by one |
| High | Proposal, contract summary, financial report, technical specification | Line-by-line comparison with the source, a second pair of eyes |
| Very high | Legal interpretation, health information, payment transaction | AI only assists; the decision and the text stay with the expert |
Method 1: Ground the model in a source
The most effective way to reduce hallucination is to make the model answer from the source you provide rather than from its own memory. You give the relevant document, table or record along with the question, and the instruction is clear: "Rely only on this text; if it is not in the text, say you do not know."
The systematic version of this approach is a setup that makes company documents searchable and hands the relevant passages to the model with every question. We explained it in detail in our article on AI that talks to your documents. Grounding in a source reduces the risk greatly but does not bring it to zero; the model can misread the source or combine two passages incorrectly.
Method 2: Ask for sources
Ask the model to show, next to every important claim, the document and section it is based on. This speeds up the reviewer's work: the answer to "Where did this come from?" is right there. A claim that cannot point to a source should be treated as unverified.
A caution: the source the model shows can also be made up. Web links and article titles given by general chat tools in particular must always be opened and checked.
Method 3: Do not leave calculations to the model
Language models are good with text and unreliable with arithmetic. Totals, averages, percentages and due-date calculations should be left to a calculation tool, a spreadsheet or a direct database query. The model is used to interpret and explain the result.
This separation is especially important in finance and ERP reports. Figures should come straight from the source system, while AI is limited to understanding the question and explaining the result in plain language. That is the core principle of the setups we build on the ERP reporting side.
Method 4: Make human approval part of the process
"Check the output" is easy to say, but it gets skipped on a busy day. Tie the check to a process, not a habit:
- AI output is never published or sent directly; it lands as a draft first.
- It is clear who approves it.
- The approver knows what to check in the output (figures, names, dates, commitments).
- Errors found are logged; recurring error types are fixed by improving instructions or sources.
Checklist
- Tasks have been classified by risk level.
- The model works grounded in your own sources wherever possible.
- Sources are requested for important claims.
- Figures come from a calculation tool or the system.
- The approval step and the approver are defined.
- Error logs are kept and reviewed regularly.
How we do it at Globya
In the AI solutions we build, we treat fact-checking as part of the design, not a warning added afterward. We ground the model in company sources, pull figures directly from the system and tie important outputs to an approval step. When needed, we build this as custom software integrated with your own systems. To clarify what you need, you can reach us through the contact page.
Frequently asked questions
Don't newer, more powerful models stop hallucinating?
Error rates may drop as models improve, but today no language model can be said to have eliminated hallucination entirely. The need for checking remains.
Does it help to tell the model "don't make things up"?
Instructions help; in particular, "if you don't know, say you don't know" and "rely only on the text provided" reduce the risk. On their own, they are not a guarantee.
Are there tools that catch hallucinations automatically?
There are methods that compare the output with the source or check it with a second model, and they are useful. Even so, for high-risk work we recommend that the final check stays with a person.
The Globya assistant is online 24/7; it answers right away and passes your question to the team if needed.