Anthropic acknowledged a massive biosafety failure: 133 million AI dialogues remained without filters.

Anthropic developers have revealed a troubling detail in their security infrastructure: for nearly a year, Claude's biological defense system was effectively inactive regarding a significant segment of user traffic. This concerns a pool of approximately 50,000 external contractors and a colossal volume—over 133 million interactions with models. This information is contained in my new analytical risk report, published in August 2026.
System failure lasting 11 months
The crux of the problem is that from May 2025 to April 2026, blocking biological classifiers were not applied to traffic from feedback collection platforms. Through this infrastructure, external contractors not only evaluated responses but also conducted open dialogues with the models. As it turned out, all this traffic was marked with an internal flag that simultaneously disabled both the blocking of dangerous requests and even the logging of their triggers.
This means that potentially dangerous requests were not just unblocked—they completely fell out of subsequent monitoring systems. Moreover, the company acknowledged weak access control among contractors: until April 2026, personnel vetting procedures at many providers did not reliably filter out potential malicious actors. By estimate, infiltrating the Red Team would not have been too difficult a task.
Audit revealed only 62 suspicious dialogues
After detecting the error, Anthropic restored nearly the entire array of correspondence, except for ~1% of data on one platform where messages were not saved without pressing the send button. To analyze all requests, the Claude Sonnet 5 model was used, which assigned a high level of biological risk to 1,197 transcripts. Of these, 757 belonged to Anthropic's own internal teams, and 378 to intentional red teaming. Thus, only 62 external dialogues were not related to testing.
Manual review of these cases and a sample of 30 red-team dialogues did not reveal clearly dangerous use capable of significantly enhancing an attacker's ability to create biological weapons. Despite this, the company acknowledged that the incident undermines confidence in the absence of other unknown gaps in protection.
Revision of risk assessments and dependence on its own AI
This case directly impacted Anthropic's official position. The risk level for the CB-1 scenario (significant AI assistance in creating known biological threats) was retroactively revised from "very low" to "low." The current level of catastrophic harm risk is now characterized as "low but not negligible." Similarly, the risk assessment for misalignment in high-stakes situations was raised.
The report also reveals a striking degree of the company's dependence on its own models. Claude writes most of Anthropic's production code, and internal research actively uses agents based on Mythos 5 and Model 2. This creates an interesting paradox: the company assessing AI risks is simultaneously its most active corporate consumer.
My comment: This case is a vivid illustration that even leaders in the AI security industry have "blind spots." The fact that the error did not lead to real incidents is more luck than merit. Investors and regulators should take this as a signal: auditing protective systems must be continuous, not retrospective. Anthropic's transparency on this matter deserves respect, but it also underscores how fragile biosecurity can be in an era of rapid AI scaling.