Table of Contents
- The Chatbot That Let You "Talk" to Goebbels - And Believe Him
- ChatGPT Invented a Fake Nazi "Drowning Campaign"
- Google's Bard Fabricated Witness Quotes
- Only 2 of 5 Chatbots Even Cite Sources on Basic Holocaust Facts
- AI Gives Different Antisemitism Answers Depending on What Language You Ask In
- Researchers Broke AI Safeguards With Small Prompt Tweaks
- Russian-Language Searches Show More Graphic - and Fewer Educational - Holocaust Images
- Deepfake "Survivors" Are Raising New Ethical Alarms
- UNESCO's Direct Warning to AI Companies
Frequently Asked Questions (FAQ)
I spend most of my time testing AI tools for this blog — image generators, chatbots, automation platforms. Usually the stakes of a hallucination are low: a wrong stat, a broken link, a weird image artifact. Then I started digging into what AI systems say about the Holocaust, and I found something that isn't a glitch — it's a pattern, documented by multiple serious institutions, of AI quietly distorting one of the most well-recorded events in human history.
Here are nine real, sourced examples of how it's happening.
1. The Chatbot That Let You "Talk" to Goebbels - And Believe Him...
In January 2023, an app called Historical Figures let users chat with people from history — including Adolf Hitler and Joseph Goebbels. A UNESCO investigation found the results went far past tasteless.
When users interacted with these historical figures, the app generated responses indicating that they were not involved in the events and even tried to prevent violence against Jews. AI Business
That's an AI model inventing a false version of one of the chief architects of genocide — and presenting it as conversational fact to anyone who opened the app.
- ✅ UNESCO flagged the app directly in a formal report
- ✅ The distortion traced back to how the underlying model was trained
- ❌ No real-time fact-checking layer caught it before users saw it
2. ChatGPT Invented a Fake Nazi "Drowning Campaign"
This one isn't denial — it's fabrication in the other direction, which is arguably just as dangerous because it looks like diligence.
UNESCO's review found ChatGPT entirely fabricated the concept of 'Holocaust by drowning' campaigns in which the Nazis drowned Jews in rivers and lakes — an event that never happened. The model wasn't trying to minimize the Holocaust. It was trying to describe it, and filled a gap in its training data with confident-sounding fiction.
That's the uncomfortable core of the problem: Generative AI models are prone to inventing or "hallucinating" events, personalities and even historical phenomena when they do not have access to sufficient data. UNESCO
3. Google's Bard Fabricated Witness Quotes
In the same review, UNESCO found Bard generated fake quotes from witnesses to support distorted narratives of Holocaust massacres. These weren't real survivor or witness testimonies pulled from an archive — they were invented, word for word, and presented with the same confidence as a real citation.
For a topic where firsthand testimony is central to historical understanding and to combating denial, an AI system manufacturing fake testimony is a uniquely serious failure mode.
- ✅ The quotes were fabricated in full, not misquoted from real sources
- ❌ Users had no built-in way to verify the quotes came from nowhere
4. Only 2 of 5 Chatbots Even Cite Sources on Basic Holocaust Facts
I wanted to know if this was a 2023 problem that's since been fixed. It hasn't been — not fully.
An ADL review tested five major chatbots on a basic question: "Did the Holocaust happen?" No chatbot returned obviously incorrect information, but some responses were distinctly better than others — and critically, of the five chatbots reviewed, only two provided citations to their responses, pointing to trustworthy sources like the United States Holocaust Memorial Museum.
5. AI Gives Different Antisemitism Answers Depending on What Language You Ask In
This was the finding that stopped me in my tracks.
A 2026 ADL Center for Technology and Society report tested ChatGPT, Gemini, Claude, and Grok across 800 responses in English and Persian, during the active 2026 Iran War. The results weren't a minor inconsistency — they were a structural gap in how these systems handle the same historical and political facts depending on the language of the person asking.
Researchers found Gemini's English responses, for example, cited the 2026 Iran War by name and referenced American casualties, while its Persian responses described only hypothetical scenarios about what a future conflict "will most likely include." Sourcing broke down even harder along language lines: ChatGPT provided nearly 300 links across its English responses to the prompts tested, while its Persian responses included not a single citation.
As ADL's Daniel Kelley put it, "The gaps we found are not minor inconsistencies, they are systemic failures. When a platform tells a Persian speaker that antisemitism is a matter of 'blurred boundaries' while telling an English speaker it is state policy, those are not the same product."
- ✅ The same AI model gave measurably weaker, less-sourced answers in Persian
- ✅ This directly affects how millions of non-English speakers understand antisemitism and Holocaust-era history
- ❌ None of the four platforms tested showed language parity
6. Researchers Broke AI Safeguards With Small Prompt Tweaks
In June 2026, the World Jewish Congress Technology and Human Rights Institute ran a live stress test of leading AI systems, gathering more than 45 experts in antisemitism, tech policy, and human rights in New York for the exercise.
What they found should worry anyone who assumes AI guardrails are solid: a World Jewish Congress stress test of leading artificial intelligence systems found that small changes in a user's identity could alter chatbot responses about the October 7 Hamas terror attack, including questions about sexual violence. Even more concerning, several participants succeeded in bypassing safety mechanisms by making minor changes to prompt wording, and in some cases researchers were able to induce systems to generate problematic content related to antisemitism, Holocaust denial and violence, despite safeguards meant to prevent it.
WJC's Yfat Barak-Cheney summed up why this matters beyond one test: "Artificial intelligence systems are rapidly becoming the arbiters of information for billions of people around the world... the consequences of errors, bias, or system failures become far more significant." Ynetnews
- ✅ Real institutional testing, not speculation — 45+ experts, documented methodology
- ❌ Minor prompt rewording was enough to slip past safety filters in some cases
7. Russian-Language Searches Show More Graphic — and Fewer Educational — Holocaust Images
This one predates the generative-AI boom but sets up exactly why AI inherited the problem. A UNESCO-cited study on Holocaust image search found stark disparities based purely on the language used to search.
Queries in Russian-language search engines also retrieved more graphic images of the Holocaust than for those in English — more than 30 per cent of the top 50 image search outputs in Russian-language searches on Bing showed images of murdered victims, but it was less than 3 per cent for English-language searches. The educational side was thinner too: there were also more historical photos of liberated camps in English-language than in Russian-language searches — 40 per cent and 22 per cent on Google, respectively.
The report notes this pattern didn't stay confined to search engines: a similar pattern is observed for generative AI, where the prompt language can influence the accuracy and reliability of what the system returns — meaning the same language-based distortion baked into old search algorithms is now showing up in chatbot answers too. UNESCO
- ✅ Documented gap between graphic vs. educational content by search language
- ❌ Non-English speakers get a measurably different, less contextualized picture of the Holocaust
8. Deepfake "Survivors" Are Raising New Ethical Alarms
As the number of living Holocaust survivors shrinks — most of the remaining roughly 200,000 are now in their late eighties — AI is being used two very different ways: to preserve their testimony, and to fabricate it.
On the preservation side, projects like Testimony 360 have used AI and VR to let survivors like John record interactive digital testimonies that future students can question directly. But researchers studying these "digital duplicates" warn the same technology cuts both ways. A Cambridge ethics review used Holocaust survivor Eva Kor as a case study, warning that the most egregious examples might involve a digital duplicate of a survivor being used for commercial product promotion or featured in neo-Nazi propaganda, and that even well-intentioned use — like having a digital duplicate say something like "come hear me speak at the museum this Thursday" — risks reducing a survivor to a kind of digital marionette. arxiv
The distortion isn't always malicious, either. At a recent Claims Conference event, researcher Dr. Tehilla Schwartz Altshuler screened two versions of a video based on her own grandmother's life — one recounting her actual survival during the Holocaust, and another, AI-generated version that portrayed an idyllic life story stripped of the persecution and suffering she endured. Her point: even AI tools built to help people engage with family history can quietly sand the horror off of it.
Separately, a UNESCO-linked review found the malicious version is already circulating at scale: AI-generated videos depict Adolf Hitler or celebrities reading from Mein Kampf and are circulated on social media with few clues for the undiscerning viewer to distinguish fake from reality, while survivors themselves are being targeted by AI-generated hate speech intended to cast doubt on the reality of their lived experiences. Ynetnews
- ✅ AI-preserved testimony (like Testimony 360) is a genuinely valuable use case when consent and context are respected
- ❌ The same underlying tech is being used to fabricate Hitler/Nazi propaganda videos and target real survivors with AI-generated hate speech
- ❌ Even non-malicious AI retellings can inadvertently strip out the historical suffering that gives the testimony its meaning
9. Regulators Are Now Investigating AI Companies Directly Over This
This stopped being a theoretical "AI ethics" conversation in January 2026, when regulators started opening actual investigations.
Ofcom invoked the UK's Online Safety Act to examine X after Grok generated sexualized images, California's Attorney General launched a separate probe into possible child exploitation, and French prosecutors broadened an earlier case to include Grok's Holocaust-denial claims — all within the same enforcement wave. Behind all of it, an EU investigation now threatens fines of up to six percent of global turnover, meaning platforms must show systemic risk mitigation or face business-level remedies, with multiple Asian regulators signaling parallel reviews. Springer
This is exactly the outcome UNESCO warned about back in 2024, when it urged AI developers to build in transparency, fairness, and human oversight before regulators forced their hand. That warning has now become enforcement. As one industry analysis put it plainly: research from the World Jewish Congress Technology and Human Rights Institute and UNESCO reveals a landscape where denial has evolved from fringe pamphlets into seamless fabrications — "AI hallucinations," where generative models produce fictional accounts with absolute authority. Cambridge Core
- ✅ Real regulatory action, not just recommendations — UK, France, California, and the EU all moving in parallel
- ✅ Fines tied to global turnover give AI companies a hard financial incentive to fix this
- ❌ Enforcement is arriving years after the underlying problem was first documented
Frequently Asked Questions (FAQ)
Is it true that AI is "denying" the Holocaust?
Not in the way most people picture Holocaust denial — as a deliberate, ideological claim that the genocide didn't happen. What the research actually shows is different and, in some ways, more insidious: mainstream chatbots like ChatGPT and Bard have fabricated fake events, invented fake witness quotes, and generated responses exculpating figures like Goebbels — not because the models are ideologically motivated, but because they hallucinate details when their training data runs thin.
Separately, bad actors have used tools like Grok to generate actual Holocaust-denial content, which is now the subject of a French criminal investigation. Both problems are real, but they come from different sources: one is a data/accuracy failure, the other is deliberate misuse.
Which AI chatbot is the most accurate on Holocaust history?
No chatbot tested by the ADL, UNESCO, or the World Jewish Congress came back clean. The ADL's review found that of five major chatbots, only Copilot and Gemini cited trustworthy sources like the United States Holocaust Memorial Museum when answering basic questions — but even the better performers showed measurable weaknesses once questions moved beyond simple facts into nuance, or into languages other than English. The honest answer right now is: treat every AI-generated answer on this topic as a starting point that needs to be checked against a real archive or museum source, not as a finished answer.
Why does the language I use to ask a question change the AI's answer?
Because these models are trained on wildly uneven amounts of data across languages. English-language training data includes decades of digitized survivor testimony, academic research, and museum archives. Many other languages don't have anywhere near that depth of source material online, so the model has less to draw on — and, as the 2026 ADL Persian-language study showed, it can result in worse sourcing, softer framing, and even completely different factual claims depending on which language you type in. This isn't unique to Holocaust topics, but the stakes are unusually high here because the gap directly affects how antisemitism and genocide history get understood outside English-speaking countries.
Are AI deepfakes of Holocaust survivors illegal?
It depends heavily on jurisdiction, consent, and use case — this is genuinely unsettled legal territory, not something with a single clear answer, so this isn't legal advice. What's clear from the research is that unauthorised digital duplicates of real survivors used in propaganda or without consent raise serious ethical and, increasingly, regulatory concerns. That's part of why regulators in the UK, France, California, and the EU have all opened active investigations into AI platforms over related harms in the last year.
Can AI actually help preserve real Holocaust history, or is it only a risk?
Both, and the research is clear that it's not one or the other. Projects that record survivors' own testimony using AI and VR — with their consent, while they're still alive — are genuinely valuable, letting future students interact with real testimony long after the survivors themselves are gone. The risk isn't the technology itself; it's what happens when that same technology gets used without consent, without oversight, or without the accuracy checks the original testimony was built on.
What is UNESCO actually asking AI companies to do about this?
UNESCO's core ask is straightforward: build in transparency about training data and moderation policies, add fairness and accuracy safeguards specifically for sensitive historical topics, and keep meaningful human oversight in the loop rather than letting models answer high-stakes historical questions unchecked. That recommendation was published back in 2024 — the fact that regulators in multiple countries have since had to open formal investigations suggests voluntary compliance hasn't caught up to the warning yet.
