
I have come into possession of a fresh Anthropic report that exposes a troubling incident in Claude's security system. It concerns a nearly year-long failure in the operation of biological classifiers—those very barriers designed to block attempts to use AI for creating biological weapons. The scale of the problem is striking: approximately 50,000 people and around 133 million interactions with the models were affected.
Critical vulnerability in contractor infrastructure
The essence of the incident is that from May 2025 to April 2026, traffic from feedback collection platforms, where external contractors tested the models, passed through without the blocking action of biological filters. The technical detail that particularly concerns me: the traffic was flagged with an internal-use marker that simultaneously disabled both the classifiers themselves and the system for logging their activations. This means that potentially dangerous requests were not only left unblocked but also left no traces for subsequent auditing.
Worse still, Anthropic admitted to weak access control among contractors. Until April 2026, many of them lacked reliable personnel vetting procedures, making it feasible for a malicious actor to infiltrate Red Team squads.
What the retrospective analysis revealed
After detecting the error, the company recovered nearly all data, except for about 1% of dialogues on one of the platforms. For verification, Anthropic used Claude Sonnet 5, which identified 1,197 transcripts with a high risk level. However, upon closer inspection, the picture changes: 757 of these were requests from Anthropic's own internal teams, and another 378 were deliberate red teaming. Only 62 external dialogues remained that were unrelated to testing.
Manual review of these cases found no explicit attempts to create biological weapons. The company assesses the likelihood of a significant increase in chemical-biological risk as "very unlikely," citing the small number of dangerous requests and their short-term nature. However, the very fact of such a prolonged gap in defense makes me doubt the completeness of their conclusions—if the classifiers were not working, then there may not have been enough signals for deep analysis either.
Revision of risk assessments
The incident prompted Anthropic to revise its own assessments. Under the CB-1 scenario (AI assistance in creating known biological threats), the risk for the past period was retroactively raised from "very low" to "low." The current level is characterized as "low but not negligible." Additionally, the risk assessment for misalignment in high-stakes situations has been raised from "very low" to "low."
Dependence on its own models
A curious detail in the report is the degree of AI integration into Anthropic's own operations. Claude already writes most of the code that ends up in the company's production repositories. This is a sign of technological maturity, but at the same time, it is a potential point of failure: if the model makes errors in code, the consequences could be systemic.
My verdict: this case is a vivid illustration that even leaders in the AI security industry have blind spots. Anthropic's ability to quickly detect and analyze the problem is encouraging, but the very fact of an 11-month gap in defense is a serious wake-up call for the entire industry. Regulators should take note: self-regulation in matters of AI biosecurity is not yet working as reliably as one might hope.