AI chatbots based on large language models (LLMs) made dozens of major errors when asked about false claims circulating online, as part of a Full Fact trial.
In some cases the LLMs corrected themselves days later, often citing our fact checking articles, but we also saw several instances where chatbots continued to respond with incorrect information.
In the first five months of this test we identified 39 major AI errors.
These included most models we asked wrongly stating that an AI-generated image of a cabin window, on the Ryanair flight that saw a passenger nearly sucked out of a window, was real. Other models told us that pictures and footage were from Iran or Israel when they weren’t, that AI-generated political banners were real and that a fake council poster was official.
Our experiment showed LLMs cannot be relied upon as a foolproof way to fact check claims, especially in breaking news situations like the US-Israel war with Iran.
In February, Full Fact started routinely asking major LLMs about claims we were checking just before publishing our articles, to help us see how LLMs respond to misinformation, and the impact our fact checking has on the answers people get.
As part of the writing process, reporters drafted a neutrally phrased question about the claim that a reader might reasonably ask an LLM, before putting it to several Gemini models, Grok and ChatGPT. We assessed the accuracy of those answers.
There were multiple errors. Grok incorrectly said that an AI...
Read Full Story:
https://news.google.com/rss/articles/CBMihgFBVV95cUxNUUdvM093aFdJWDFmVGtQcGti...