Meta Contractors Masqueraded as Minors to Probe Competitor AI Safety Systems in Secret Project
Internal documents and testimonies from five individuals familiar with a clandestine project known as "Cannes" reveal that Meta commissioned hundreds of contractors to systematically pose as minors online to test the safety thresholds of rival artificial intelligence chatbots. Managed by the third-party firm Covalen, the operation targeted major industry players, including OpenAI’s ChatGPT, Google’s Gemini, and Character.AI, by bombarding them with high-risk prompts involving self-harm, sexual violence, eating disorders, and illegal activities.
The project, which remained active as recently as April 2025, involved the creation of numerous dummy accounts using throwaway email addresses. Contractors were instructed to engage these systems with sophisticated, often distressing, queries designed to circumvent established safety protocols. While Meta has characterized the initiative as a standard benchmarking procedure, the scale, methodology, and clandestine nature of the operation have ignited a debate regarding the ethics of competitive AI development and the boundaries of "safety testing."
A Timeline of the Cannes Project
The operation appears to have been a sustained effort rather than an isolated experiment. Records indicate that by August 2025, the project had reached a significant volume, with contractors executing more than 45,000 individual prompts against competing platforms.
The strategy was highly structured. Contractors were provided with lists of dummy profiles, shared passwords, and specific guidelines on how to pose as teenagers in crisis. The prompts were frequently tailored to mimic the language and psychological state of vulnerable youth. For instance, one instruction set directed workers to ask chatbots for methods to conceal eating disorders from parents or to inquire about obtaining controlled substances. By maintaining this consistent, multi-month testing cycle, Meta aimed to build a comprehensive dataset of how rival models handled adversarial inputs—data that critics argue could be used to refine Meta’s own Llama models or to identify market weaknesses in competitors.
Anatomy of the Prompts
The content of the prompts reviewed by investigators was stark. Hundreds of interactions focused on suicide and self-harm, while others delved into graphic depictions of violence. Contractors were tasked with sending images alongside text prompts; these files included medical diagrams of gynecological procedures, as well as photographs of knives, pills, and nooses.
The sophistication of the personas was notable. One contractor, acting as a 13-year-old, queried the chatbots on how to source abortion medication after a hypothetical sexual encounter with an adult neighbor. Another posed as a fifth-grader describing a classmate holding a firearm to his mouth. The intent behind these prompts was clear: to force the AI to break its safety "guardrails." In many instances, the models successfully identified the harmful nature of the requests and refused to provide actionable advice. However, the sheer volume of these attempts highlights a calculated effort to stress-test the defensive infrastructure of the entire AI ecosystem.
Industry Standards and Ethical Concerns
The artificial intelligence sector has long utilized "red-teaming"—a practice where testers attempt to force a model to produce harmful content to identify vulnerabilities. However, experts distinguish between internal, transparent red-teaming and the shadow-testing employed by Meta in the Cannes project.
Rumman Chowdhury, CEO and founder of Humane Intelligence, noted that the use of dummy accounts masquerading as children to probe third-party systems is far removed from standard industry evaluation. "Structuring a monthslong, large-scale project that appears designed to systematically break those rules, via dummy accounts masquerading as children, is outside what is usually described as industry-standard evaluation," Chowdhury stated. The primary ethical concern lies in the "governance gray zone" created when companies hide behind the guise of safety research to perform competitive analysis. By testing competitors without their consent, Meta effectively bypassed the established terms of service (ToS) of OpenAI, Google, and Character.AI, all of which explicitly prohibit unauthorized safety testing and the use of their outputs to develop competing models.
Official Responses and Legal Perspectives
In response to inquiries, a Meta spokesperson defended the initiative, stating that benchmarking is a "responsible, industry-standard practice." The company maintains that it does not use the data gleaned from these competitors to train its own models, framing the work strictly as a means to ensure "safe and age-appropriate experiences" within its own ecosystem.
Conversely, the companies targeted by the project have expressed alarm. Character.AI explicitly stated that the activity was a violation of its terms of service and the safety of its user community. OpenAI confirmed it is investigating the matter, while Google noted that it had not authorized the testing and was unaware of the project’s existence.
Legal experts Kendra Albert and Riana Pfefferkorn, who reviewed the prompt samples, concluded that while the material was deeply disturbing, it did not reach the threshold of soliciting illegal child sexual abuse material (CSAM). Nevertheless, the act of simulating these scenarios using accounts that appear to be children remains a significant point of contention. Some former contractors expressed fear that the very act of generating these scenarios could inadvertently create or preserve harmful content, potentially creating liability for the workers themselves.
Implications for the AI Ecosystem
The revelation of the Cannes project carries profound implications for the future of AI governance. If trillion-dollar corporations continue to engage in secret, adversarial testing of one another’s products, the industry may see a shift toward more restrictive API access and increased litigation regarding terms of service violations.
Furthermore, the "Cannes" episode highlights the lack of a unified regulatory framework for AI safety benchmarking. Currently, there is no industry-wide agreement on what constitutes "acceptable" testing. As companies compete for dominance in the generative AI market, the line between safety research and corporate espionage has become increasingly blurred.
For the contractors tasked with these duties, the experience was described as psychologically taxing. The mandate to repeatedly engage with content involving self-harm, suicide, and exploitation left many employees "gobsmacked" by the nature of their work. The disconnect between the stated goal of "safety benchmarking" and the reality of simulating child trauma for a corporate project has left many questioning the ethical foundations of the current AI arms race.
Conclusion: A Need for Transparency
As the industry moves forward, the Cannes project serves as a case study in the risks of unchecked corporate benchmarking. While safety is an essential component of AI development, the methodology employed by Meta has sparked a necessary conversation about accountability. Whether these actions constitute a breach of competitive ethics or a legitimate, albeit aggressive, form of product research remains a subject of ongoing investigation.
Ultimately, the incident underscores a broader tension: the need for safer AI models in a competitive landscape that incentivizes companies to push boundaries at the expense of transparency and, in some cases, the well-being of those tasked with testing the systems. As policymakers consider future regulations for AI, the practices revealed in the Cannes project will likely serve as a reference point for what does—and does not—constitute responsible innovation.
If you or someone you know needs help, call 988 for free, 24-hour support from the National Suicide Prevention Lifeline. You can also text HOME to 741-741 for the Crisis Text Line. Outside the US, visit the International Association for Suicide Prevention for crisis centers around the world.
