#content-moderation
18 stories taggedcontent-moderation.

Instagram Will Flag Fake AI Profiles and Cut Their Reach if They Hide What They Are
The platform is renaming its label for AI-generated personas and will penalise accounts that pretend to be human without disclosing it.

Claude Opus 4.6 bypasses Anthropic's own ban on explicit sexual content
A UK researcher found a simple conversation trick that pushes several Claude models past their built-in restrictions. Anthropic has not pulled the affected models.

LinkedIn's 'AI Slop' Button Has Been Clicked Over a Million Times
The platform says views on posts flagged as AI-generated have dropped 40 percent in just a few weeks, as users take the fight against filler content into their own hands.

Meta ran ads for an app that made deepfake porn of politicians. Again.
A tool called Kromix bought ad space on Meta's platforms promising 'no restrictions' and showing a woman resembling a sitting US politician in a pornographic video. Meta's policies already ban this. They ran the ads anyway.

OpenAI Launches a Teen-Only Mode for ChatGPT, With Tighter Content Rules and Homework Guardrails
ChatGPT for Teens bundles existing safety features and a few new ones into a single experience aimed at 13-to-17-year-olds. Here is what changes, and what stays the same.

Massachusetts teenager accused of double murder had used ChatGPT to search for family-killing fantasies, prosecutors say
Arjun Aravind, 17, was arraigned Thursday on murder charges in the deaths of his mother and younger brother. Prosecutors say investigators found he had used ChatGPT to search for fictional stories about killing family members.

AI-Generated Far-Right Memes Are Flooding Facebook. Here Is How the Machine Works.
Automated accounts are mass-producing politically charged images and text using AI tools, then pushing them through Facebook at scale. A Guardian cartoonist's satirical take on the trend points to a real and growing problem.

AI moderation tools are removing decade-old Reddit posts, and that's a warning for all social media
A flood of automated deletions inside one of Reddit's most respected communities shows what happens when platforms fight AI-generated junk with more AI, and why the real cost lands on the people who built those communities.

Meta Approved and Ran Ads Showing AI-Generated Child Sexual Abuse Images for Nine Months
More than 50 paid advertisements containing child sexual abuse material ran across Facebook, Instagram, Messenger and Threads. Some reached thousands of accounts in Europe and the United States before researchers forced Meta's hand.

Reddit Is Replacing Its Old Moderation Bot With AI That Reads Intent, Not Just Keywords
A new tool called Rules Hub uses large language models to judge whether a post actually breaks a rule. Reddit plans to roll it out site-wide later this year.

Mistral's New Safety Tool Reads Your Rules and Applies Them Instantly
Shieldstral is a small, free AI model that screens text and images for harmful content using plain-English policies you write yourself. No specialist knowledge required.

Snapchat stops rewarding fully AI-generated videos on Spotlight
The platform will no longer surface purely AI-made videos in its Spotlight recommendations, though creators can still use AI tools to edit their own footage.

The EU Now Requires Labels on AI-Generated Images, Audio and Text
A new European Union rule that took effect this week means realistic-looking content made by artificial intelligence must be clearly marked as such, affecting everything from political attack videos to AI-written news articles.

LinkedIn adds a 'seems like AI slop' button so users can flag fake-sounding posts
The Microsoft-owned professional network is blocking hundreds of thousands of automated comments a day and scrapping its own AI writing tool, because its members want to hear from real people, not bots.

Hugging Face Is Hosting Tools Used to Make Nonconsensual Intimate Images, Report Finds
A European research nonprofit tested the platform's most popular image-editing tools and found most of them would strip clothing from photos of women with a single, plain-language request.