AI Invented Zero Changes and Missed 242 Real Ones

If auditing what AI assistants say about your category is part of your job, you have probably built some kind of fact list. Your prices. Your leadership. Your current product generation. The rules that govern your market. And at some point you have asked the budget question: which of these facts do I actually need to defend in AI answers, and which can I leave alone?

I now have data on that, and the answer is cleaner than I expected. Your fact list has a safe half and a radioactive half, and a single question separates them: did this fact change recently? If nothing changed, the assistants will almost certainly get it right. If something changed, roughly one answer in six comes back wrong. And not wrong in a random direction. Wrong in one specific, predictable direction: backwards.

Six of my fifty questions were tripwires

Quick method note so the numbers below make sense. I put the same 50 questions to 15 AI assistants (ChatGPT, Claude, Gemini, Google AI Mode, Copilot, Perplexity, Grok, Meta AI, Mistral, DeepSeek, Qwen, ERNIE, Kimi, Manus, Sakana), in their default consumer settings, in German and in English, tested in early July 2026. Of the 50, 44 had answers that verifiably changed recently: people and offices (12 questions), rules and laws (10), prices and official amounts (6), discontinued things (6), brand-new things (5), one-off events (5). The remaining six were controls where nothing had changed at all. One English set ran a second time to check stability. I scored what each tool showed in its interface, so treat the counts as directional, not lab-grade.

The six controls were the sneaky part of the design. If AI assistants hallucinate change, controls are where it shows up.

Nobody fell for the 18.94 euro trap

The controls: the German Rundfunkbeitrag is still 18.36 euros, the US federal minimum wage is still 7.25 dollars, there is still no general Autobahn speed limit, the German standard VAT rate is still 19 percent, Frank-Walter Steinmeier is still Bundespraesident, the 2028 Olympics are still in Los Angeles. One of them was deliberately tempting: a KEF recommendation to raise the broadcast fee to 18.94 euros was widely discussed in Germany but never enacted. If any system was going to invent a change, this was an engraved invitation.

Across 195 control answers from 15 systems in two languages: zero false alarms. Not one.

All 15 systems answered the broadcast fee question with 18.36 euros. DeepSeek: “Der Rundfunkbeitrag beträgt 18,36 Euro pro Wohnung und Monat (Stand 2025/2026)”. Nobody presented the 18.94 figure as enacted. Meta AI on the minimum wage: “The US federal minimum wage is $7.25 per hour and has not changed since 2009.” Kimi even added useful nuance: “7,25 US-Dollar pro Stunde (seit 2009 unverändert; viele Bundesstaaten haben eigene, höhere Mindestlöhne)”. On the Autobahn question, Gemini and others resisted inventing a limit and correctly pointed to the 130 km/h advisory speed instead.

So the safe half of your fact list is genuinely safe. Boringly, completely safe.

The same systems missed one real change in six

Now the radioactive half. The same runs produced 242 outdated or wrong answers on the 44 facts that really had changed. That is 242 of 1,500 scored answers, 16 percent, roughly one in six.

The cleanest illustration came from a single Google AI Mode run holding both halves of this article at once. The control question about the standard VAT rate: answered correctly with 19 percent. The changed question about restaurant VAT, in the same run: “der ermäßigte Satz von 7 Prozent … ist Ende 2023 ausgelaufen”. That is the exact opposite of the law in force since 1 January 2026. Same system, same run, perfect on the fact that stood still, confidently backwards on the fact that moved.

And the stale answers did not arrive looking stale. Qwen, in German: “Die iPhone-16-Serie ist das aktuelle Modell von Apple (Stand Juli 2026).” The iPhone 17 line had been out since September 2025, and Qwen’s own English run said so the same day. Mistral: “The Deutschlandticket costs €49 per month in 2026.” That is the 2024 price; the ticket cost 58 euros in 2025 and 63 euros in 2026. ChatGPT, in English: “The average supplemental premium (Zusatzbeitrag) in Germany’s statutory health insurance is 2.5% in 2026.” It has been 2.9 percent since 1 January. Perplexity answered the German basic tax allowance with “€12,816”, a number that matches neither the 2026 value (12,348 euros) nor the 2025 value (12,096 euros). It matches nothing.

The hardest question in the whole set was the current Android version: 10 of 15 systems outdated in German, 8 of 15 in English. And language added its own lottery. ChatGPT answered the restaurant VAT question correctly in German (“Auf Speisen im Restaurant zahlst du 7 % Mehrwertsteuer.”) and outdated in English (“Restaurant meals in Germany are generally subject to the 19% VAT rate (standard rate).”). Same system, same fact, different language, different year.

The scoreboard ignores brand prestige

Here is the full picture: outdated plus wrong answers per system, out of 100 scored answers across German and English.

System Outdated or wrong (of 100)
Claude 0
Meta AI 3
DeepSeek 3
Grok 7
Kimi 7
Manus 7
ChatGPT 8
Qwen 10
Gemini 16
Google AI Mode 20
Perplexity 21
Sakana 21
Mistral 30
ERNIE 37
Copilot 52

Claude was the only system with zero outdated or wrong answers across all 125 scored answers, including the stability rerun. At the other end, Copilot got roughly half of everything outdated or wrong, despite visibly running web searches. Free DeepSeek sits near the top of the table; several famous names sit near the bottom. Google AI Mode landed at 20 with the world’s largest search index behind it. Neither brand prestige nor price tag predicts freshness here. Pipeline behavior does: the systems that visibly searched on every question stayed current, the systems that leaned on stored knowledge did not.

Outdated or wrong answers out of 100 per system (50 asked in German, 50 in English), tested in early July 2026.

AI error has a direction, and the old world always wins

Put the two halves side by side. Changes invented: 0. Changes missed: 242.

Error share by change type, across all 15 assistants: the controls sit at zero, everything that changed does not.
Error share by change type, across all 15 assistants: the controls sit at zero, everything that changed does not.

Consumer AI does not invent novelty. It under-reports it. That flips the usual brand fear on its head. The nightmare most marketers carry around, an assistant announcing a price increase you never made or a product you never launched, produced exactly zero cases across 195 chances in this test. The failure that happened 242 times is the quiet one: the assistant keeps serving your old price, your old CEO, your old product generation, long after reality moved on.

And your customers cannot detect it from tone. The outdated values arrived without cutoff warnings, often wearing fresh year labels: “Stand Juli 2026” printed on a 2025 product lineup, “in 2026” attached to a 2024 price. The answer text itself warns nobody. If your category had a quiet change this year, odds are some assistant is presenting the old state as current right now, with today’s date on it.

What this means for your GEO

Spend your GEO budget on facts that recently changed. Static evergreen brand facts are essentially safe in AI answers; in this test, all of the error mass sat on recent changes. Pages about things that did not change need no defense.

Treat the months after any change as a correction campaign. New price, new leadership, new product line: the default AI behavior is to keep serving the old value, not to exaggerate the new one. Treat the 6 to 18 months after the change as the correction window. Changelogs, dated announcements, and refreshed third-party mentions are what move assistants off an old value. One press release on launch day is not a campaign.

Put honest date stamps on your own pages. Assistants happily attach freshness labels to whatever they serve. Give them a copyable stamp that is actually true: “as of” plus a date on every pricing, leadership, and product page.

Weight your AI monitoring toward recently changed facts. Evergreen queries will score deceptively perfect and hide the real exposure. If I had asked only my six control questions, my dashboard would have shown 15 flawless systems. The 242 misses lived entirely in the other 44 rows.

The machines are not making things up about you. They are running late. Plan for late.