The Anthropic team has disclosed data on a critical gap in the biological safety system of its flagship model, Claude. The failure, which persisted for nearly a year, affected about 50,000 external contractors and touched approximately 133 million interactions with the artificial intelligence. This information is contained in my new analytical report on the state of protective mechanisms in the industry.

System Failure: 11 Months Without Blocking Classifiers

For 11 months — from May 2025 to April 2026 — all traffic passing through feedback collection platforms was processed without the activation of blocking biological classifiers. These systems are designed to detect and prevent attempts to use the model for potentially dangerous chemical or biological purposes. The critical error was that traffic was marked with an internal flag that simultaneously disabled the filters and disabled logging of their activations.

The most alarming aspect is that most participants had the ability to conduct open dialogues with the models, rather than only evaluating ready-made responses. This means that potentially dangerous requests were not only unblocked but also completely fell out of view of subsequent audit systems. Moreover, the report acknowledges that before the tightening of requirements, many contractors lacked reliable personnel screening procedures, making it possible for malicious actors to infiltrate Red Team squads.

Scope of Review: 1197 Transcripts Under Suspicion

After detecting the error, Anthropic managed to recover nearly the entire dataset, except for 1% of conversations on one platform, where information was not saved without pressing the submit button. For a retrospective analysis of all requests, the Claude Sonnet 5 model was deployed, which assigned a high risk level to 1197 transcripts.

A detailed breakdown showed that most of them (757) belonged to Anthropic's own internal teams. Of the remaining 440 cases, 378 were deliberate red teaming — specialized attempts to break the defenses to test their reliability. Thus, only 62 external dialogues were not related to testing. Manual review of these conversations and a random sample of 30 red-team dialogues did not reveal clearly dangerous use that could significantly enhance malicious actors' potential to create biological weapons.

Despite the company assessing the likelihood of a substantial increase in chemical-biological risk as "very unlikely," the very fact of discovering such a large-scale vulnerability forced a revision of official assessments. The risk level for the CB-1 scenario (creation of known threats using AI) was retroactively raised from "very low" to "low." A similar adjustment also affected the risk assessment for misalignment in high-stakes situations.

It is worth noting that the report also reveals the extent of Anthropic's own dependence on its AI systems: Claude already writes a large portion of the company's production code. However, from my point of view, this incident underscores a fundamental problem across the entire industry: even leading laboratories with multi-layered defenses may have "blind spots" in their infrastructure that they are unaware of. This is a serious wake-up call for the entire market, demonstrating that AI safety is not a static set of filters, but a continuous process of auditing and revisiting one's own assumptions.