Strategy May 28, 2026

Folie à Deux in the Machine: When Models Share Hallucinations

Riddles leave models with nowhere to hide

Károly Boczka
DeepSeek and Grok hallucinations

This is a guest blog authored by Humane Intelligence and its volunteersexploring topics related to AI evaluations and sociotechnical topics in AI.


In February 2026,
Anthropic publicly confirmed something the AI community had long suspected: DeepSeek conducted industrial-scale distillation from Claude outputs, using 24,000 fraudulent accounts and 16 million exchanges to systematically extract capabilities and train their own models. The announcement sent shockwaves through the AI governance community. 

In October 2025, while running the  Hungarian Riddles Benchmark for the first time, we found that DeepSeek and Grok produced nearly identical wrong answers on 26 out of 100 Hungarian cultural riddles, not just failing on the same questions, but fabricating the same content, often in the same wording.

Click on the screenshot for the whole sheet of common hallucinations.

The pattern made the finding impossible to dismiss. Unlike open-ended cultural questions where a model can bluff its way to a plausible answer, a riddle has one correct solution rooted in cultural knowledge. Youeither know that a specific Budapest landmark is guarded by a tongueless lion, or you don’t. No paraphrasing saves you. Riddles are a perfect stress test: short, culturally non-negotiable, and binary. They are deterministic by nature. There is one correct answer, while language models are fundamentally non-determanistic . That tension is exactly what makes riddles a hard and revealing task. There is nowhere for a model to hide. 

The phenomenon has a human equivalent. Not officially recognized as a standalone psychiatric diagnosis today, but widely referenced in clinical literature. In psychiatry, folie à deux describes a condition where two individuals who share the same environment or prior beliefs develop identical delusional thinking, reinforcing each other’s false reality. The parallel is striking: language models trained on overlapping data can converge on the same false content, each reinforcing the other’s fabricated world.

This type of content-level, cross-model hallucination overlap had not been systematically documented in academic literature or public forums. There is no public documentation to be found about what the dataset shows. Only fragments point in the same direction:

DeepSeek’s own technical documentation explicitly states: “In line with Grok-1, we have evaluated the model’s mathematical capabilities using the Hungarian National High School Exam.” This is a rare public acknowledgment of shared methodology between the two models, and it involves Hungarian, the same language where the convergent hallucinations appeared. It concerns mathematics, not cultural knowledge, but the overlaps may not end here.

A widely-circulated Reddit thread asked: “Is Grok-3 just DeepSeek R1 in disguise?” The author, testing both models with identical prompts in Russian, observed the same speech patterns, the same response logic, and the same quirk of inserting Chinese characters randomly into answers. It mentions behavioral convergence between the same two models but no identical hallucinated content was documented. 

Researchers at Salesforce AI identified what they call a “shared imagination space,” the phenomenon where models trained on overlapping data converge on similar fabrications when confronted with questions they cannot answer from knowledge. It is close but only theoretical, without concrete examples of what those fabrications actually look like across models.

When retesting both models in February 2026, the pattern had dissolved. The newer versions of Grok and DeepSeek solved several of the 26 previously failed riddles correctly. The remaining failures were no longer shared: each model now hallucinated differently, independently. The convergence had disappeared as silently as it had appeared. Nothing was explained publicly.  

This is a governance gap. If two models can silently converge on the same fabricated cultural content and then silently diverge without explanation, the standard tools of AI oversight are insufficient. Benchmark scores measure aggregate performance, not content-level hallucination overlap. Model cards report general capabilities, not behavioral lineage. Transparency frameworks do not yet require disclosure of training relationships that could explain why two nominally independent models think alike. For low-resource languages, this gap is wider: less public data means fewer external checks, fewer native researchers tracking failures, and fewer mechanisms to detect model failures.

We know models can absorb each other’s outputs at industrial scale. But take one example from the illustration above: ‘Meg és a Mogorva’ meaning ‘Meg and the Grumpy One’. Two characters that both models identified as from a Hungarian cartoon, but they do not exist anywhere in Hungarian or any other culture. The closest real reference is ‘Meg and Mog,’ a British children’s book about a witch and her cat. But it is a corrupted echo, not a source. Two models fabricated the same non-existent character independently. Can shared training data alone explain that? The honest answer is: we do not know

And yet, the geopolitical framing around these two models could not be more different: one built under the banner of American innovation, the other routinely cited as a national security concern to the United States from the other side of the world. If the threat narrative is real about one model, what does that mean for converging hallucinations with another model? Shared training data alone may not produce identical fabrications. For that, there would need to be shared architecture, shared methodology, shared learning patterns, possibly a shared upstream source, or deliberate distillation. All of these possibilities deserve scrutiny.

These are precisely the transparency gaps that governance needs to address. What we do know is that low-resource languages are both exposed and can be the first to expose the problem. Like a canary in a coal mine, they surface failures that richer-resource languages can obscure for longer. As argued in our previous post, multilingual models cannot be assumed to be culturally neutral. This finding extends that argument. Models may not only fail differently across languages, they may fail identically, for structural reasons invisible to standard evaluation. That is not a benchmark curiosity; it could be a warning sign of systemic coupling between supposedly independent models.

 

Sign up for our newsletter
Sign up for our newsletter