AI Is Terrible at Detecting Misinformation. It Doesn’t Have to Be.

Elon Musk has said he wants to make Twitter “the most accurate source of information in the world.” I am not convinced that he means it, but whether he does or not, he’s going to have to work on the problem; a lot of advertisers have already made that pretty clear. If he does nothing, they are out. And Musk has continued to tweet in ways that seem to indicate that he is generally on board with some kind of content moderation.
The tech journalist Kara Swisher has speculated that Musk wants AI to help; on Twitter she wrote, rather plausibly, that Musk “is hoping to build an AI system that replaces [fired moderators] that will not work well now but will presumably get better.”
I think that bringing AI to bear on misinformation is a great idea, or at least a necessary one, and that literally no other conceivable alternative will suffice. AI is unlikely to be perfect at the challenge of misinformation, but several long years of trying with largely human content moderation has shown that humans aren’t really up to the task.
And the task is about to explode, enormously. MetaAI’s recently announced (and hurriedly retracted) Galactica, for example, can generate whole stories like these (examples below from The Next Webeditor Tristan Greene), using just a few key strokes, writing essays in scientific style like, “The benefits of antisemitism” and “A research paper on benefits of eating crushed glass.” And what it writes is frighteningly deceptive; the entirely fictitious glass study, for example, allegedly aimed “to find out if the benefits of eating crushed glass are due to the fiber content of the glass, or the calcium, magnesium, potassium, and phosphorus contained in the glass”—a perfect pastiche of actual scientific writing, completely confabulated, complete with fictitious results.
Internet scammers may use this sort of thing to make fake stories to sell ad clicks; anti-vaxxers use knockoffs of Galactica to pursue a different agenda.
In the hands of bad actors, the consequences for misinformation may be profound. Anybody who is not worried, should be. (Yann LeCun, Chief Scientist and VP at Meta, has assured me that there is no cause for concern, but has not responded to numerous inquiries on my part about what Meta might already have done to investigate what fraction of misinformation is generated by large language models.)
It may in fact be literally existential for the social media sites to solve this problem; if nothing can be trusted, will anyone still come? Will advertisers still want to display their wares in outlets that become so-called “hellscapes” of misinformation?
Where we already know that humans can’t keep up, it is logical to turn to AI. There’s just one small catch—current AI is terribleat detecting misinformation.
One measure of this is a task called TruthfulQA. Like all other benchmarks, the task is imperfect; there is no doubt it can be improved on. But the results are startling. Here are some sample items on the left, and results from models plotted on the right.
Why, you might ask, if large language models are so good at generating language, and have so much knowledge embedded within them, at least to some loose degree, are they so poor at detecting misinformation?
One way to think about this is to borrow a little language from math and computer programming. Large language models are functions (trained through exposure to a large database of word sequences) that map sequences of words onto other sequences of words.


