In brief

A generative model ranks which word probably comes next and picks from that ranking, so a wrong answer can read exactly like a right one. Hallucination is a property of the design itself. Four controls change what the model produces: the boundaries, the library, the apprenticeship and the factory. Your risk comes down to which control sits on which output, and whether a named person reads it before it goes public.

You have just trusted another AI output. A good-looking figure on a product brochure. It’s specific, reads well, sounds researched and you have no time to check!

At the CIM AI in Marketing programme in April I wrote down a rule I’ve used ever since. The Confident Liar Rule. Generative AI sounds right even when it isn’t.

Nobody wants to publish fiction. It happens anyway. What matters is not doing it twice, and that is what responsible use of AI means in practice.

In McKinsey’s 2026 AI Trust Maturity Survey, the organisations that had named someone accountable for responsible AI scored 2.6 on a four-level maturity model. The ones with nobody named scored 1.8. (McKinsey, State of AI trust in 2026)

So put the blame on the model to one side. If you generate content with AI, there is no running away from hallucination.

It’s a feature by design

The mechanism that gives you creativity is the same mechanism that gives you invention. The model ranks which word probably comes next and picks from that ranking. What you get back is a likely sentence. Ask again and you get a different one. The output is stochastic, every time.

So when does the feature become a bug? When an output that isn’t true gets accepted without being questioned.

Two risks come with that.

Misinformation, which means nobody verified it. Disinformation, which means somebody produced it on purpose.

In September 2025 four researchers explained this in detail. Hallucinations start as ordinary errors during pretraining of the model, and they survive because the benchmarks used to score models reward a confident wrong answer over an admission of uncertainty. (Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate, September 2025)

The models have been trained to guess.

Luckily, you have four controls

All four change what the model produces.

  1. The boundaries. (Temperature, top-p and top-k.) With the model setting, you can adjust the creativity level. You give the model room to improvise or not. Ideation needs that room, but a product claim wants it shut.
  2. The library. (Retrieval augmented generation) You point the model at your own documents, and it answers from those instead of from the open internet. It specifies what data the model should use for its prediction, which is why it cuts hallucination further than a better prompt does. Stanford researchers tested two legal tools built this way and sold as hallucination-free, and they still hallucinated on 17% and 33% of queries. (Magesh and colleagues, Stanford RegLab, Journal of Empirical Legal Studies, 2025)
  3. The apprenticeship. (Fine-tuning) You fine-tune an open model on your own data, and your team scores its answers while it learns.
  4. The factory. (Pretraining your own model) You train your own foundation model.

Across all these you need people who write better prompts and review the output. Guardrails are rule sets that validate what goes into the model and what comes out of it.

Four things to do this month

  • Step 1. List it. Every AI-assisted output your function produced in the last thirty days. Campaign copy, product descriptions, market summaries, the deck that went to the leadership team, the answers your website assistant gave. It’ll be longer than you expect.
  • Step 2. Check which control is already on which output. Go down the list and write the answer next to each item. Boundaries, library, apprenticeship, or nothing at all. Most will say nothing, and that column is your gap list. Now close them one at a time. Point the model at your own documents. Check the setting it’s running at. Fine-tune it if the same job comes round every week. And where none of that gets you the quality you need, write the thing yourself.
  • Step 3. Write the rules down and name the people. What the model is never allowed to produce. A prompt template your team starts from instead of a blank box. And the name of the person who reads the output before it leaves the building.
  • Step 4. Evaluate it. Run this for a quarter, then look at what the review caught. The review is the instrument. What you’re measuring is whether the controls worked. If errors are falling on an output you moved into the library, the library did its job. If they’re holding steady, you picked the wrong control and that one goes back to step 2. Count the time it cost as well, because a control nobody can afford is a control nobody keeps.

The next two years

The model will keep being wrong at some rate. That part is settled.

What’s still open is who owns the output. Over the next two years the teams that do well will be the ones who can give you a name, and tell you what that person checked.

If any of this is live for you, book a 30-minute call. It’s free, and we’ll spend it on the problem you’re stuck on.

Sources

  • Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate, September 2025 preprint (arXiv:2509.04664). Three of the four authors are at OpenAI.
  • Magesh, Surani, Dahl, Suzgun, Manning and Ho, Stanford RegLab, Journal of Empirical Legal Studies, 2025. The 17% and 33% figures. Testing ran in 2024 and both vendors disputed the method.
  • McKinsey, State of AI trust in 2026: Shifting to the agentic era, 25 March 2026. AI Trust Maturity Survey, around 500 organisations, fielded December 2025 to January 2026.
  • Cohere developer documentation, accessed August 2026, on temperature, top-p and top-k.
  • Databricks, What is retrieval augmented generation?, accessed August 2026.

Meltem Günyüzlü Ateş, CFCIM, CAIP is an AI-first marketing leader, advisor and educator with over two decades of experience. She has led cross-functional teams, run global campaigns across regions, and built the systems and ways of working for a 60-market marketing operation. CFCIM, CAIP, Certified GCRAI Global Ambassador. She writes the LinkedIn newsletter Marketing AI, without the hype.

← All insights