Automated Moderation is Here to Stay; Accountability Must Keep Pace
EFF Deeplinks
- Automated content moderation systems often reproduce existing biases, fail to grasp context, and disproportionately harm journalists, activists, and marginalized communities.
- Meta and other platforms have demonstrated high error rates in moderation, particularly regarding non-violent speech in certain languages and the misclassification of protected content.
- Accountability must evolve alongside AI; companies and policymakers should adopt the Santa Clara Principles 2.0 to ensure human rights and due process are integrated into moderation.
Core Problems
- Linguistic Inconsistency: Models struggle with low-resource languages (e.g., Maghrebi Arabic, Kiswahili) due to poor data labeling and limited native speaker involvement.
- Overzealous Moderation: Automated systems frequently fail to distinguish between harmful content and legitimate expression, resulting in systemic suppression.
- Vulnerable Communities: Repeated misclassification of LGBTQ+ content as explicit or the suppression of political speech, such as content related to Palestine.
Eight Recommendations for Policy and Practice
- Augmentation over Replacement: AI should prioritize content for human review rather than replace human judgment.
- Transparency: Platforms must be clear about if and how automation influences content decisions.
- Auditing for Bias: Regular assessments focusing on low-resource languages, marginalized groups, and conflict zones.
- Right to Appeal: Users must have accessible paths to contest automated decisions, with human moderators handling these appeals.
- Human Rights Impact Assessments: Platforms should conduct and publish regular evaluations of their moderation practices.
- Third-Party Accountability: Strict adherence to these principles must be required of all third-party moderation vendors.
- Regulatory Caution: Legislators should avoid mandates that require automated moderation or dictate specific technical design choices regarding expression.
- Multi-Stakeholder Oversight: Decisions around moderation design must include input from civil society, researchers, and affected communities.