AI safety claims need controls across every access path
Anthropic processed 133 million contractor interactions for roughly a year while biological-risk filters were disabled and warnings were not logged.
The control gap covered traffic from vendors providing human feedback and approximately 50,000 contractors between May 2025 and April 2026. The crucial distinction is that 133 million counts total traffic, not harmful requests. Anthropic reported that its retrospective review found no clear evidence of misuse. However, the review relied largely on the company’s own classifier and limited manual checks.
AI-safety claims should be supported by evidence that controls are actually working on every access path and that they produce auditable records. Contractor platforms belong within that system boundary too.