Meta said on Wednesday that more than 97% of the content it acted on was found by its own systems before users reported it. In India, the company acted on 5.3 million pieces of child sexual exploitation content during the same period, with more than 98% detected before user reports.
The new tools include a large language model (LLM) system to detect what Meta calls “signposting” — ads that may appear harmless but are suspected of directing users to illegal content or other harmful activity elsewhere online. Meta says this is a tactic it has recently seen bad actors use as they change their methods to avoid detection. While the ads themselves may not contain illegal material, they can act as a gateway to websites hosting child sexual abuse material.
Meta said it is now looking at where an ad sends users, not just what the ad contains. The company can use this information to block websites or other destinations that break its rules and take action against the accounts behind them.
The company is also using additional AI-driven scans to find child exploitation content that earlier systems may have missed, and it said it will continue adding new signals as it learns more about how these networks operate.
Another new tool is a “red-teaming AI agent” that tests Meta’s own safety measures. It looks for weaknesses that bad actors could use to circumvent the company’s protections. Meta said this could help it find new methods of abuse before they become more common.
Additionally, Meta is improving its systems for detecting people who return to its platforms with new accounts after their previous accounts have been removed.
The measures come as Meta continues to face pressure over child safety on its platforms. The company has faced lawsuits and criticism from lawmakers. In August, Meta reached an agreement to pay up to $18 billion to settle a child safety lawsuit involving 29 U.S. states.
Meta has introduced several child-safety features this year, including parental controls for Meta AI, pre-teen accounts on WhatsApp, alerts for parents when children search Instagram for self-harm content, and additional parental controls for teen accounts on WhatsApp.