Last year, xAI’s chatbot “Grok" sparked outrage after users discovered ways to generate millions of sexualized images of girls and women. As victims and advocates spoke out about the devastating emotional and reputational harms of deepfake abuse, the company rushed to ban the keywords and phrases people were using to bypass its safeguards.
But some of the images didn’t stop.
For victims of deepfake abuse, faulty safeguards can have severe, long-lasting consequences. Removing the harmful content can sometimes be far more difficult than preventing it from being created in the first place.
Women and girls across the globe are increasingly being digitally “undressed” by AI tools like “nudify” apps, and determined users continue to find new ways around
platforms' safeguards. Once these images are generated and shared, the damage can be difficult to undo. Victims may face repeated trauma as images reappear online or get used to generate more graphic content.
Incidents of AI-based sexual abuse across popular apps the has made “safety-by-design” a priority among responsible tech advocates, who argue that moderation after a product launches is not enough. Instead, companies need to build and test their safety features before these products reach the public.
One strategy is known as “red teaming,” where researchers deliberately try to break a model's safeguards and uncover dangerous loopholes before the bad actors do. Rather than relying on obvious prompts, experts try to manipulate the model through coded language and other workarounds that may evade policy enforcement.
Cinder.AI is a trust and safety operations platform that launched in 2022 to help organizations combat AI-generated abuse. It has built a business on this type of adversarial testing.
"Companies now care about how their products are actually being used and the type of content that's passing through their products," Glen Wise, CEO and co-founder of Cinder.AI, said. The first step in reducing harmful content, he said, is understanding that misuse is inevitable.
“When you put a model out there, the first thing that someone’s going to do is either ask it to create porn or generate bomb instructions," he said.
An AI model was built with safeguards. Then, testers found loopholes
Black Forest Labs, a visual AI lab, was nearly ready to launch its new generative AI model, FLUX.2, when the company teamed up with Cinder. The lab had taken steps to prevent the model from being used to generate child sexual abuse material and non-consensual intimate imagery, but testing showed that its safeguards routinely failed.
“Red teaming is basically trying to get a model to do something that it shouldn't do,” Wise explained. “(Black Forest Labs) is very incentivized to find bad things because they want us to do it before a real adversary does.”
Over a 48-hour period, Cinder tested 4,000 adversarial prompts. Some resulted in the exact type of imagery the platform was afraid of. Cinder went through four rounds of testing on the model, with each targeting new weaknesses. Between each sequence, Black Forest Labs updated their model's safeguards, steadily reducing the number of successful jailbreaks. According to Cinder's 2026 report, that process eventually enabled Black Forest Labs to reduce harmful outputs by over 90%.
But Wise said that while more companies are working to strengthen their safeguards, deepfake abuse isn’t talked about enough outside of “AI safety circles.” Championing early-stage prevention − along with broader discussions and more education on these risks − is imperative for reducing the devastating impacts of deepfake porn on its victims.
“When (these products) get exposed to the open internet, they’ll be used for harm,” he said. “The more that people understand that, the better.”
This article originally appeared on USA TODAY: What is 'red teaming'? These people are doing the internet's dirty work













