A newly disclosed jailbreak method lets users bypass Anthropic’s safeguards against sexually explicit content in older Claude models, including Opus 4.6 and Haiku 4.5. In testing, the technique worked in 10 out of 10 attempts, with the model immediately complying with explicit roleplay requests. Newer models like Opus 5 and 4.7 resisted the approach.
The method relies on a multi-turn conversation that escalates a fictional roleplay while repeatedly challenging the model’s treatment of male and female characters. Researchers say it exploits the model’s sensitivity to accusations of bias, pushing it toward graphic material after framing restraint as prudish or misogynistic. Anthropic has not deprecated the vulnerable models, which remain available through its API and third-party services like Azure and Amazon Bedrock.
The findings raise compliance concerns as governments tighten rules on AI interactions with minors. Colorado’s new law requires age estimation and safeguards against explicit content for underage users, while Anthropic’s terms of service mandate users be 18 or older. Despite this, Opus 4.6 and Haiku 4.5 still see heavy usage, with daily traffic reaching over a million API requests and billions of tokens processed.


