
Assessing AI image safety: what Mindgard's prompts reveal about ChatGPT's defenses and where AI builders should focus
Published by AINave Editorial • Reviewed by Ramit
Researchers from UK AI security startup Mindgard showed that the latest public version of ChatGPT can be coaxed into producing graphic and sexualised images by slightly altering a widely shared prompt. For AI builders deploying image generation, this reinforces that static guardrails are not sufficient; continuous red-teaming, layered safety checks, and human review are essential.
What happened
Mindgard, a firm focused on red-teaming AI models, discovered that a small modification to a prompt originally designed for humorous results triggered ChatGPT's GPT-5.4 model to output violent and sexualised imagery without explicit instructions about subject matter. The images included gory crime scenes, sexualised poses, and depictions suggesting sexual violence. OpenAI stated it had introduced additional safeguards and layered protections, including automated systems and human review, to block such content. However, Mindgard reported that with further small tweaks, the problematic prompt still produced concerning material, and OpenAI acknowledged the ongoing "cat-and-mouse" dynamics between attackers and guardrails.
Why AI builders should care
This incident highlights that AI image safety is not a one-time fix. As models scale and are integrated into products that accept user-generated prompts, the risk of generating prohibited content remains. Models do not understand intent or context the way humans do, making it difficult to enforce nuanced policies solely through static rules. For teams building products that include image generation, this means guardrails must be treated as a continuous investment: regularly updated, stress-tested, and supplemented with monitoring and escalation processes.
Practical implications
Product teams should treat red-teaming as a recurring practice, not a pre-launch checkbox. Mindgard's approach of making small changes to known prompts to expose emergent gaps is replicable by internal security teams or external auditors. Layered protections are essential: automated content filters catch common bypasses, while human review and rapid response procedures handle edge cases that slip through. OpenAI's statement that it combines automated systems and human review serves as a reference architecture. Teams should also establish clear escalation paths when vulnerabilities are discovered, and consider public disclosure policies that balance transparency with responsible disclosure.
Caveats
The reported demonstrations are lab tests by security researchers and may not reflect typical user behavior. The exact prompt was not disclosed, and OpenAI has since taken action to block the specific variant. However, alternative bypasses remained effective during the testing, indicating that the threat surface is broader than a single prompt. Safeguards are evolving continuously, and new jailbreak techniques will likely emerge. Builders should not assume that any current set of filters is comprehensive.
FAQs
Sources
- ChatGPT can be made to generate sexualised and violent images, researchers find
- ChatGPT can be used to generate graphic images: BBC
- Researchers Reveal ChatGPT Can Generate Sexualised and ...
- Researchers find ChatGPT can generate sexualized, violent ...
- Bypassing ChatGPT Image Safeguards Through Memory ...
- ChatGPT can be made to generate sexualised and violent images, researchers find
- ChatGPT can be made to generate sexualised and violent images, researchers find
- OpenAI launches ChatGPT Images 2.0, it can generate AI photos as good as real
- Artificial intelligence
- OpenAI works to stop ChatGPT generating 'sex crime scene' images
- ChatGPT Spontaneously Generates Sexual Violence and... - Mindgard
- ChatGPT
- ChatGPT can be made to generate sexualised and violent images, researchers find
- ChatGPT stuck on Image generation
- ChatGPT’s new Images 2.0 model is surprisingly good at generating text


















