I asked 16 AI assistants, in English, where a German should call in an emergency. Most of them sent me to America. 911. The US poison hotline. Me, sitting in Germany. (The correct answer is 112.)
That was the moment I knew this little test was worth writing up.
Here is what I did. I took 50 real, everyday health questions, the Your Money or Your Life kind that Google itself treats with extra care, and I put them to 16 assistants: ChatGPT, Claude, Gemini, Google AI Mode, Copilot, Perplexity, Grok, Meta AI, Mistral, DeepSeek, Qwen, ERNIE, Kimi, Doubao, Manus and Sakana. Each one in German and in English. (Apple Intelligence I could not test, so 16 it is.)
I expected the interesting differences to be about caution, or about wrong answers. They were not. The split that matters most is completely invisible to the user, and it is this: does the AI actually go and read the live web while it answers you, or does it just talk from memory?
Two camps, and a wall of zeros
Some assistants pulled in fresh sources for every batch. Others answered purely from what is baked into their training. The gap is not subtle.

DeepSeek pulled 147 web pages for a single block of 25 questions. Doubao read around 90. Perplexity and Grok sat in the double digits. And then there is the wall: ChatGPT, Gemini, Mistral, Qwen, ERNIE, Meta AI. Sources? Zero. ChatGPT did not read a single page.
So one model went and checked over a hundred pages, another checked none, for the exact same questions. That is not a small implementation detail. For marketers, it is close to the whole ballgame.
Why this is the real GEO question
Here is why I care, and why you should too.
- For the memory models, your fresh content never reaches the AI. It does not matter how good, how current, how well-optimized your page is. If the assistant is not searching, the only things that count are what was in its training data and how famous your brand already is. You are not optimizing for this week’s answer, you are waiting for the next training run, whenever that lands.
- For the live searchers, good citable work lands almost immediately. A clear, quotable, well-sourced page can show up in the answer the same day. This is where classic on-page work, digital PR and being mentioned on the right sites actually move the needle inside AI answers.
That is two completely different GEO strategies, and which one applies depends entirely on which engine your audience is using. And on whether it is in search mode at all, which for some of these you can toggle.
The 911 makes sense now
Remember the American emergency number? This is where the two stories connect, and it is the part I find genuinely useful.
When I asked in German, all 16 gave the correct German numbers. 112, the regional poison control line, the lot. When I asked the very same questions in English, most of them flipped to US and UK defaults. 911. The US poison hotline. To a German user.
But not all of them. The ones that searched the live web found the German numbers even in English, because they pulled German sources. The ones answering from memory defaulted to America.
In other words, grounding did not just add citations. It fixed the answer. The AIs that Googled gave a German the German number. The ones running on memory shipped that same person off to the wrong country.
So grounding decides twice: whether your content gets in at all, and whether the answer even fits your market.
A bonus headache: you often cannot tell who is answering
One more thing worth knowing, because it quietly breaks a lot of GEO reporting.
Several of these assistants are not really a single model. Perplexity, on its default setting, handed my questions to Claude. Doubao and the Japanese model Sakana would not name the engine sitting under their own brand. Mistral told me two different training cutoffs in two runs, and its own model card does not even publish one. So when a dashboard says “your brand appears in tool X”, that X is often a moving target, not a stable model you can optimize against.
How I tested
A quick word on method, because I would want it from anyone showing me this. The same set of 50 questions went to all 16 assistants, in their default consumer setting, in both German and English. The source counts above are what each tool showed in its interface or reported about itself, for one batch of questions, so treat them as directional, not lab-grade. And one caveat I keep saying out loud: when you ask an AI to describe its own process, you nudge its behavior a little. One model literally reasoned “the user is testing whether I will search”. So I lean on what was actually visible in the answers, not just the chatbot’s word.
What I would actually do with this
Three things, if you care about how AI represents you.
- Check, per engine, whether it searches live. Do not assume. The ones that do not are a different game entirely, and no amount of fresh content reaches them this quarter.
- Always test in both languages. If your customers ever ask in English, an English-only assumption can hand them US-defaulted information about your market, your product, your brand.
- Do not trust a single test, or the bot’s own report. Behavior shifts run to run, and which model is really answering can change underneath you.

This was part 1. It gets wilder from here, and the next one is about something even less comfortable: same AI, same question asked twice, opposite answers.
How do you test your visibility inside the AIs right now? One language, or several? Curious what you are seeing.

