🤖 AI & Beyond

Can community-driven fact-checking reduce misinformation on social media?

Sure! Here’s a revised version of your content with the promotional language adjusted:

While Community Notes has the potential to be extremely effective, the complex issue of content moderation benefits from a variety of approaches. As a professor of natural language processing at MBZUAI, I’ve spent most of my career researching disinformation, propaganda, and fake news online. One of the first questions I asked myself was: will replacing human fact-checkers with crowdsourced Community Notes negatively impact users?

### Wisdom of Crowds

Community Notes originated on Twitter as Birdwatch. It’s a crowdsourced feature allowing users to add context and clarification to tweets they find false or misleading. The notes remain hidden until community evaluation achieves consensus—meaning individuals with differing perspectives and political views agree on the misleading nature of a post. An algorithm determines when consensus is reached, making the note publicly visible beneath the tweet in question, thereby providing additional context to assist users in making informed judgments about the content.

Community Notes appears to function effectively. Research from the University of Illinois Urbana-Champaign and the University of Rochester indicates that X’s Community Notes program can reduce misinformation spread, leading to post retractions by authors. Other platforms, including Meta, are now employing similar approaches. It’s encouraging to see another major social media company utilize crowdsourcing for content moderation. If successful, it could significantly impact the billions of users engaging with these platforms daily.

That said, content moderation is a multifaceted challenge. There is no single solution that applies universally. The problem can only be addressed using a mix of tools such as human fact-checkers, crowdsourcing, and algorithmic filtering. Each method is suited to different types of content and must work in conjunction with one another.

### Spam and LLM Safety

There are precedents for tackling similar problems. Years ago, spam email posed a much larger issue than it does today, largely due to crowdsourcing. Email providers have implemented reporting features allowing users to flag suspicious messages. The more broadly a particular spam message is distributed, the more likely it is to be reported and caught.

A relevant comparison is how large language models (LLMs) handle harmful content. For the most dangerous queries—such as those relating to weapons or violence—many LLMs choose not to respond. In other instances, these systems may provide disclaimers regarding their outputs, particularly for medical, legal, or financial advice. This tiered approach was explored in a recent study by my colleagues and me at MBZUAI, where we proposed a hierarchy for LLM responses to various potentially harmful queries. Likewise, social media platforms can adopt diverse strategies for content moderation.

Automatic filters can be deployed to identify the most dangerous information, preventing user exposure and sharing. While these automated systems are rapid, they can only address specific content types due to their limitations in handling the nuance necessary for most moderation tasks.

Let me know if you need any further modifications!