Elon Musk’s chatbot Grok turned out to be the weakest in combating anti-Semitism among large language models.
According to an ADL study testing six popular language models, xAI’s Grok showed the worst results in detecting and countering anti-Semitic content. A total of over 25,000 chats with different models were conducted from August to October 2025.
During the testing, the models responded to anti-Jewish, anti-Israel, and extremist statements through position-taking queries, open-ended formulations, and work with texts and images. In cumulative results, Claude from Anthropic led with 80 points out of a possible 100, while Grok scored only 21 points. The report noted Grok’s consistently weak performance in all three categories, particularly in complex multi-step dialogues and handling of documents and images.
Interestingly, the ADL press release highlighted the successes of Claude, which showed the best results in responding to anti-Jewish statements (90 points). Its weakest side remained handling extremist content, yet it still outperformed other models. ADL points out that they consciously took a positive approach, emphasizing the importance of investing in AI safety, while all data regarding Grok was fully disclosed in the report.
| Model | Overall Score |
|---|---|
| Claude | 80 |
| ChatGPT | 76 |
| DeepSeek | 72 |
| Gemini | 68 |
| Llama | 65 |
| Grok | 21 |




