Crypto news

17.08.2026
09:58

Anthropic acknowledged a massive failure in Claude's biosafety systems: 133 million dialogues remained without filters.

ii-startap-Anthropic-AI

Anthropic developers have disclosed a serious incident in the biological safety system of their Claude models. The failure, which lasted nearly a year, affected a pool of approximately 50,000 external contractors and about 133 million AI interactions. This information became key in the new Risk Report published by the company.

Systemic Failure: 11 Months Without Blocking Classifiers

For 11 months—from May 2025 to April 2026—all traffic on the platforms through which third-party specialists worked with the models was processed without activating blocking biological classifiers. These systems are designed to intercept potentially dangerous requests in the fields of chemistry and biology. As it turned out, the traffic passed with an internal flag that simultaneously disabled both the blocking and the logging of protective mechanism activations.

This means that potentially harmful requests were not only not rejected but also did not enter the subsequent audit system. Moreover, the company admitted weak access control among contractors: before stricter requirements were introduced, many providers lacked reliable personnel vetting procedures, making it possible for infiltration into Red Team units.

Scope of Review: 1,197 Suspicious Dialogues

After detecting the error, Anthropic restored nearly all correspondence data, except for only 1% of traffic on one platform where information was not saved without pressing the send button. To analyze the datasets, the company used the Claude Sonnet 5 model, which identified 1,197 transcripts with a high level of risk.

The distribution proved revealing: 757 dialogues belonged to Anthropic's internal teams, and 378 were deliberate red teaming—attempts by specialists to provoke the model into dangerous responses. In the end, only 62 external correspondences were not related to testing. Manual review of these cases and a random sample of red-team dialogues did not reveal any obvious dangerous use capable of significantly enhancing an attacker in creating biological weapons.

Revision of Risk Assessments and Dependence on Its Own AI

This incident forced Anthropic to retrospectively revise the risk assessment for the CB-1 scenario (using AI to create known biological threats) from "very low" to "low." The company characterizes the current level of catastrophic harm risk as "low, but not negligible." The risk assessment for misalignment in high-stakes situations was also raised.

Notably, the report also reveals Anthropic's deep dependence on its own models: Claude writes a large portion of the code accepted into the company's production repositories. Although this accelerates development, it is still less than twofold.

My comment: This case is a striking illustration that even leaders in the AI security industry have "blind spots." The fact that the failure lasted nearly a year and was only discovered retrospectively raises serious questions about the reliability of monitoring systems. For the market, this is a signal: the audit of protective circuits must be continuous and multi-layered, rather than relying on one-off checks. Investors and users should be more attentive to companies' assurances about the security of their platforms.