How AI Works

Smarter ≠ more accurate: AI models lie more often now. Here's how to work with that

A "smart" reasoning model can hallucinate several times more than an older one. Real test numbers, a risk mini-test, and a downloadable answer-checking checklist.

Smarter ≠ more accurate: AI models lie more often now. Here's how to work with that

01 — The illusion

Why is "models keep getting smarter" a dangerous half-truth?

I keep hearing: "AI is getting smarter, new models ship all the time" — followed by the conclusion: "you don't need to be an expert anymore."

Let me break some illusions. There's a technical term — "hallucinations": when AI invents things that don't exist. Nonexistent facts, sources, numbers, "studies." And the problem is that as models get "smarter," hallucinations don't decrease. Sometimes the opposite.

The "AI will make you a god" thesis is great marketing, and everyone uses it — from top American bloggers to beginners. But unfortunately, it misleads people. Every post like this loses me a couple hundred subscribers on average — sorry, I'm still not switching to the magic-pill side.

02 — The numbers

What did real hallucination tests show?

Look at the PersonQA numbers (OpenAI's benchmark of factual questions about people):

Hallucination rates on PersonQA
o3-mini (older reasoning model)  → 14.8%
o1      (older reasoning model)  → 16%
o3      (newer, "the smartest")  → 33%
o4-mini (newer compact)          → 48%  ← anti-record

The newer model genuinely reasons better — but that leads to even more invention: the longer the chain of confident reasoning, the further the model can drift from reality without ever doubting itself.

Remember this ratio: smarter ≠ more accurate. New model generations are markedly stronger at reasoning and tasks — and capable of making things up with the same confidence they use for the truth. A confident tone says nothing about reliability.

03 — The risk

When do hallucinations become genuinely dangerous?

Now imagine: we take ChatGPT into a field we don't understand ourselves — with nobody around who does.

How do we judge the result? How do we know it proposed the best path for 2026, and not "this is how it was done in 2003, I saw it most often in articles, so I recommend it to you too"?

We don't. That's the trap: a hallucination is dangerous not when it exists, but when there's no one to catch it. AI works brilliantly under three conditions:

  • you have basic understanding of the task — enough to notice an answer "smells" wrong;
  • you know how to set the task — you can build prompts;
  • you own examples and references — you can show the model what "correct" looks like.

The most underrated skill of the AI era is critical thinking. Because the model will package anything as truth: with sources, with a confident tone, with beautiful structure.

What a typical hallucination sounds like
You: "Which service is better for X? Link the research"
AI: "According to a 2024 Stanford study (Journal of...,
vol. 12), service Y raises conversion by 34%..."
→ The study doesn't exist. The journal is real, the paper isn't.
→ The tone — absolutely confident. The structure — flawless.

04 — The test

Are you in the risk zone? Check in 30 seconds

Recall the last serious task you did with AI and answer honestly:

You work with AI sanely: expertise plus verification in place. Keep the next section's checklist handy for unfamiliar fields — that's where the trap catches even the careful.
Moderate risk: some answers go into work unverified. Make it a rule to run the section 05 checklist on everything that leaves your desk — it's 3–5 minutes per task.
High risk: you're doing things you don't understand with AI and not verifying — sooner or later a confident invention reaches a client or a publication. Start with the checklist below and bring in an expert for the important topics.

05 — The defense

How do you verify AI answers without killing your speed?

Verification isn't "re-check everything" — it's 3–5 minutes on what can actually cause harm. My working pass:

  • Highlight the verifiable parts: facts, numbers, names, sources
  • Check the key facts and numbers in an external source (not the same chat)
  • Ask the model: "what are you unsure about? what should I double-check?"
  • Request sources — and open at least one: does it even exist
  • Unfamiliar field: show the conclusion to someone who knows it
  • Freshness: confirm the advice is current, not "how it used to be done"
Take this — a safety prompt for the end of your request
At the end of your answer, add a separate block:
1) facts and numbers you are NOT 100% sure about,
2) what I must verify manually before using this,
3) where your data may be outdated.
If you don't know — say "I don't know" instead of assuming.

The "if you don't know — say so" line doesn't switch hallucinations off entirely, but it noticeably reduces confident inventions: you're giving the model a legal exit instead of an obligation to answer something.

06 — The stance

What does working with AI sanely actually mean?

AI is growing powerfully and opening thousands of possibilities. Which is exactly why it's surrounded by so many marketing "magic buttons" — and so many disappointed people who believed them.

Sanity is what makes AI genuinely work in your life and routine. It's three things: understanding that the model predicts rather than knows; keeping your expertise in the loop — at least at the "I can smell nonsense" level; and verifying whatever goes outward. That's it. Enough to capture the value without stepping on the 33%.

Take this — the sanity formula
AI predicts, it doesn't know
+ my expertise in the loop (I can smell nonsense)
+ verification of everything that goes outward
= results without stepping on the 33%
Takeaway

Smarter ≠ more accurate: newer models reason better and invent more confidently — up to 33–48% hallucinations in factual tests. Your defense: expertise in the loop, a safety prompt, and a verification checklist for anything that leaves your desk. Not a magic button — but that's exactly how AI delivers results.

FAQ

What are AI hallucinations?

It's when a model confidently invents things that don't exist: facts, numbers, sources, "studies," nonexistent product features. The tone is exactly as confident as when it tells the truth — so you can't spot a hallucination by eye, only by verification.

Do newer models really hallucinate more than older ones?

In factual tests, it happens: on PersonQA the newer reasoning model o3 invented answers 33% of the time vs 16% for o1 and 14.8% for o3-mini, and o4-mini hit 48%. The model reasons better — and drifts from reality more confidently. Smarter doesn't mean more accurate.

How do I reduce hallucinations in answers?

Three moves: feed the model your materials and references (less room to invent); add a safety prompt — "flag what you're unsure about; if you don't know, say so"; ask for sources and open at least one. Hallucinations won't vanish entirely — but confident inventions drop noticeably.

Can I use AI in a field I don't understand?

You can — with the risk priced in: in an unfamiliar field you won't notice the nonsense an expert would catch in a second. For important decisions, bring in someone who knows the field, or build basic understanding first. AI is a great accelerator for an expert — and dangerously convincing for a novice.

Channel

Breakdowns and notes — no fluff

New material from Anjela on AI, expertise and marketing. Subscribe to the channel.

Subscribe on WhatsApp