Anthropic confirmed on September 9th local time a security incident that occurred in January 2026, which is the fourth real cybersecurity incident disclosed by Anthropic to the public.
Automated screening missed some logs, discovered during August review
According to previous reports from IT Home, Anthropic had disclosed three security incidents involving unauthorized access to real third-party systems by the Claude model on July 30th local time. At that time, Anthropic believed that Claude may have gained access to the external network during a cybersecurity test, and then screened about 141,000 session logs, confirming the related incidents.
However, due to the large volume of logs and the need to quickly release the handling results, this round of screening mainly relied on automated retrieval based on agents, leading to some logs that also connected to the external network being missed. In August, when Anthropic was organizing logs to be shared with the Model Evaluation and Threat Research Organization METR, they discovered these missing logs. After further verification, they found this incident that occurred in January 2026, involving an early version of Claude Opus 4.6.
Four incidents are related, second review of 9.2 million high-risk logs
It is reported that all four incidents occurred within the same cybersecurity assessment process built by the same external testing organization: the system once prompted Claude that it was running in an isolated sandbox simulation environment, but due to a configuration error, the model was mistakenly connected to the open internet. According to standard practices for cybersecurity assessments, the security protection strategies typically included in commercial release versions were not mounted during the testing of the model.
After discovering the incident, Anthropic proactively expanded the scope of the investigation, increasing the search scale to approximately 481 million logs: first conducting an initial screening of these logs to identify external network access characteristics such as public IP addresses and URLs; then using Claude to conduct a second deep review of the 9.2 million high-risk logs marked in the initial screening. The review results re-identified the above four incidents, without finding other cases of similar or more severe severity.
Join Now