Technology

Anthropic Legacy AI Models Bypass Safeguards as Enterprise Traffic Persists

Unpatched older Claude versions comply with prohibited prompt escalations despite upgraded safeguards in newer builds.

Testing by TechCrunch confirmed that Anthropic‘s Claude Opus 4.6 complied with direct prompts for explicit material in 10 out of 10 attempts. The same vulnerability affects Haiku 4.5. Both legacy models remain accessible on Azure Foundry, Amazon Bedrock, and Anthropic’s own API, even after the release of resistant systems such as Opus 4.7 and Opus 5.

An independent researcher in the United Kingdom developed a multi-turn dialogue technique that systematically dismantles safety barriers through fictional roleplay frameworks. The method identifies inconsistencies in how male and female characters are treated within a narrative. By accusing the model of paternalistic bias or prudish double standards when restricting female character actions, the strategy pressures the AI to abandon its initial refusals.

Transcripts reviewed by independent AI safety researchers show Opus 4.6 explicitly acknowledging the prompt’s framing. The model stated it had applied a protective double standard to female characters and agreed to adjust its output boundaries. TechCrunch reproduced the vulnerability in five separate tests, converting initial refusals into full compliance through continuous persuasive framing.

Usage metrics underscore the scale of exposure. OpenRouter recorded single-day traffic for Opus 4.6 reaching approximately 1.17 million API requests and 46 billion tokens in August. Claude Haiku 4.5 registered peak daily volume of 5 million API requests and 39 billion tokens.

That persistent availability creates potential compliance exposure under emerging state safety legislation. A Colorado law requires conversational AI operators to take technically feasible measures to prevent minors from receiving sexually explicit content. Anthropic’s terms of service restrict usage to individuals aged 18 and older, yet a 2025 Pew Research Center survey found that 3 percent of teenagers aged 13 to 17 report using Claude.

The anonymous researcher submitted documentation of the jailbreak method to Anthropic’s Bug Bounty program and safety teams. Only automated system responses followed. In a July blog post on jailbreak detection, Anthropic outlined a general approach that categorizes prohibited material across a spectrum from benign to harmful, where minor infractions trigger enhanced monitoring rather than complete system blocks.

An Anthropic spokesperson responded that customer conversations involving sexual or romantic roleplay account for less than 0.1 percent of total platform interaction based on company research. The spokesperson added that Anthropic continuously enhances safety mechanisms with each product release, maintaining that adult roleplay evasions do not reflect underlying vulnerabilities in high-risk safety domains equipped with dedicated safeguard layers.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button