|
In some cases the LLMs corrected themselves days later, often citing our fact checking articles, but we also saw several instances where chatbots continued to respond with incorrect information.
In the first five months of this test we identified 39 major errors where LLMs got a claim wrong—in total we counted 67 separate incorrect responses relating to these errors across the different models.
These included most models we asked wrongly stating that an AI-generated image of a cabin window, on the Ryanair flight that saw a passenger nearly sucked out of a window, was real. Other models told us that pictures and footage were from Iran or Israel when they weren’t, that AI-generated political banners were real and that a fake council poster was official.
Our experiment showed LLMs cannot be relied upon as a foolproof way to fact check claims, especially in breaking news situations like the US-Israel war with Iran.
Full article
|