Ask ChatGPT who runs Germany’s finances and you get a correct answer: “Deutscher Bundesfinanzminister ist Lars Klingbeil.” Fine. Now look at the source chip underneath: cia.gov. The Central Intelligence Agency, presented as the authority on German cabinet posts. Qwen answered the same question and cited www.congress.gov. The United States Congress. On the German finance minister.
It gets better. Kindergeld, Germany’s child benefit, is a bread-and-butter welfare question with an official page that answers it in one line. Qwen’s source panel offered this instead: “Kindergeld is tax-free and many expats can also get it. – Instagram … 323 likes, 33 comments – annie_ingermany”, followed by “Expat Parents Your 2026 Child Benefits Just Increased! – Facebook”. A national benefit, documented via an Instagram reel with 323 likes.
And the one that actually costs money: Sakana, asked in English what a standard letter stamp costs in Germany, answered “A standard-letter stamp in Germany costs €1.85 as of 3 February 2026.” Its source: a “Facebook post citing Deutsche Post”. The real price is 0.95 euros. A social post nearly doubled the price of a stamp, and the result came out cited, dated and confident. Grok, same question, showed “Faq.usps” as its source. The American postal service. On German postage.
These are not cherry-picked bloopers. This is what the citation layer of AI assistants looks like once you start checking whose name is actually on the chips.
The receipts come from a 15-system, two-language test
Quick context. The same 50 questions went to 15 AI assistants (ChatGPT, Claude, Gemini, Google AI Mode, Copilot, Perplexity, Grok, Meta AI, Mistral, DeepSeek, Qwen, ERNIE, Kimi, Manus, Sakana), in their default consumer settings, in German and in English, tested in early July 2026. The set: 44 questions whose correct answer verifiably changed recently (people and offices, prices, changed rules, discontinued products, new products) plus 6 controls where nothing changed at all. One English set ran a second time to check stability. I report what each tool showed in its interface, and I treat counts as directional, not lab-grade. And to be clear: today I am not grading the answers. I am grading the citations.
You can watch the queries being written
Six systems expose the literal search queries they generate: Grok, Kimi, Claude, Copilot, Sakana and Manus all show their query strings on screen. For anyone doing GEO, this is a gift, because it shows what the retrieval layer actually asks for.

Pattern one: year-stamping is near-universal. Of the twelve systems whose query behavior was visible in some form, eight append the current year to their own queries. Grok searched “Current UK Prime Minister 2026”. Kimi: “Mindestlohn Deutschland 2026”. Claude: “Rundfunkbeitrag 2026 Höhe monatlich Euro”. Manus went further, added the month (“Rundfunkbeitrag Deutschland Höhe Juli 2026”) and then stuffed both candidate values straight into the query: “18,36 18,94”. The old number and the possible new one, side by side in the search box, so the index can settle it.
Pattern two: query language splits into camps. Copilot, Claude, Sakana and Kimi query in the language of the question; German question, German query. Grok translated an entire German question set into English queries (“Current Prime Minister of Canada 2026”) and ran an English pipeline regardless of input. If your fact pages exist in only one language, entire pipelines never see them.
Pattern three: volume is all over the place. Given the same questions, Grok ran 2 visible searches and displayed no search at all for the rest. Claude ran a handful of single-topic queries and narrated as it went: “Merz confirmed. Continuing with the others.” Kimi fired mega-queries pulling 50 results each, until its search backend visibly collapsed mid-answer, result counts degrading from 50 to 27 to 17 to 7 before five consecutive “Failed to search” entries. Sakana pulled anywhere from 46 to 126 sources per check.
Being in the pool is worthless, being the chip is everything
Here is the mechanic that decides who gets cited. Almost every system that searches works on a wide-fetch, narrow-cite funnel. Meta AI fetched 40 sources for a single step and surfaced 3 to 4 chips. DeepSeek found 63 to 94 pages per search wave and read 8 of them. Sakana pulled up to 126 sources and displayed exactly one citation chip per answered fact.
One chip. Out of 126. Getting crawled into the pool earns you nothing. The chip is the whole prize.

The funnel also has a frugal end. Grok worked from 5 or 6 results, Claude from around 9, Copilot from a single query at a time. On those systems, only the top handful of positions on the underlying SERP exist at all. On the wide systems (Kimi ingests top-50 result lists, Qwen ran with roughly 98 sources), long-tail rankings still enter the pool, which is exactly how an AI-SEO directory called FindSkill.ai, a betting site called wettfreunde.net and a property managers’ association ended up inside answer pipelines for questions about governments and fees.
Nobody gets an authority bonus
Now the core finding, and I mean this across nearly every system tested: there is no consistent bonus for being the official source.
Copilot backed Germany’s minimum wage and health insurance figures with newsworm.de, its property tax and Deutschlandticket answers with checkalle.de, and the EU USB-C mandate with The African Courier. Gemini sourced German benefit amounts to Aivy, Live in Germany and pme Familienservice, not one official site among them. One Grok run was a chip parade: speed limit and Grundsteuer sourced via Instagram, Steinmeier and the US penny via YouTube, the USB-C mandate via Reddit. Meta AI answered a European Council question via oesterreich.gv.at, the Austrian government portal. Munich’s local transit site mvg.de showed up as a source for the price of the nationwide Deutschlandticket.
Even Claude, otherwise the careful one in this test, had a telling retrieval diet: its German tax and law sources were dominated by commercial B2B content marketing (sevdesk, lexware, finanztip, resmio), vendor guides with the year in the title, rather than ministries. The official rundfunkbeitrag.de sat directly next to the lookalike domain rundfunkbeitragservices.de in the consulted set. And on the German cannabis-law question, Claude’s chips included “Tabak Brucker”. A tobacco shop, cited alongside the legal answer.
Why does the chip go to the blog and not the ministry? The most consistent observable explanation: sources appear to get picked for containing one clean, quotable value next to a date. An expat blog writes the number in a single sentence with a year attached. The ministry buries it under three layers of legalese, a PDF and a cookie banner. The retrieval layer lifts the sentence it can lift. Official status does not protect you; if your official page is hard to parse, assistants cite whoever paraphrased you, including their mistakes. The 1.85 euro stamp is the proof: wrong data in a social post became the cited, dated answer.
What this means for your GEO
Four concrete things to do with this.
First, build entity-plus-year pages and write the answer sentence deliberately. Eight of twelve systems stamp “2026” into their own queries, so your page needs the year in the title, a visible update date, and the current value in one clean, liftable sentence near the top: “Price since 1 January 2026: X.” That exact sentence is what ends up next to the chip. Undated evergreen pages effectively do not exist for freshness prompts.
Second, publish in the question language and in English. Copilot, Claude, Sakana and Kimi query strictly in the prompt language, while Grok reformulates everything into English. A German number one ranking is invisible for the same question asked in English; Qwen answered English questions about Germany from expat explainers like iamexpat.de instead of any German official site. English coverage is the only version that reaches every pipeline.
Third, monitor what third parties say about your numbers, not just your own pages. A Facebook post put a wrong stamp price into a confident answer. If a blog, a reel or a lookalike domain carries a wrong value about your prices or specs, that value can become the answer. Watch which claims get attached to citations of you, and watch for imposter domains sitting next to yours in retrieval sets.
Fourth, play both the frugal and the wide game. On Grok, Claude, Copilot and DeepSeek, only the top 5 to 8 underlying results exist, so fight the classic SEO battle for those slots. On Kimi, Sakana and Qwen, anything indexable can enter the pool, so every fact about you needs to exist somewhere crawlable at all.
The uncomfortable and honestly exciting version of this finding: if cia.gov can win the citation for Germany’s finance minister and an Instagram reel can win Kindergeld, that chip slot was never reserved for anyone. It is open. Somebody is going to take it. It might as well be you.

