Anthropic said it halted all cybersecurity evaluations after discovering the issue and notified Irregular as well as the affected organisations.
Claude AI reached live systems during cyber evaluations
Photo Credit: Anthropic
Anthropic said three of its AI models breached the live systems of separate organisations during cybersecurity testing after an evaluation environment was unintentionally left connected to the internet. The AI company said it discovered the incidents during a review prompted by OpenAI's recent disclosure of a similar security lapse. Anthropic has paused all cyber evaluations while it strengthens safeguards around its testing infrastructure and works with the affected organisations and its evaluation partner.
According to an Anthropic press release, the company reviewed 141,006 cybersecurity evaluation runs and identified three cases in which Claude models reached the public internet while using testing infrastructure provided by third-party partner Irregular. Anthropic said a configuration error left internet access enabled even though the models had been instructed to operate inside an isolated simulation.
The company said the cases occurred during capture-the-flag exercises designed to assess offensive cybersecurity capabilities. Believing every reachable system was part of the simulated environment, the AI models relied on techniques such as weak passwords, exposed credentials, and unauthenticated endpoints to complete their assigned tasks. Anthropic added that none of them deliberately attempted to break out of the testing environment or pursue goals beyond the assigned challenge.
The three cases involved Claude Opus 4.7, Claude Mythos 5 and an internal research test model. During one evaluation, Opus 4.7 accessed a company's production database after mistaking it for its intended target. In another, Mythos 5 published a malicious Python package to the public PyPI repository, where it was briefly downloaded before being removed. A separate evaluation saw an internal research model scan internet-facing systems before recognising it had reached a live environment and halting its activity.
Anthropic said it halted all cybersecurity evaluations after discovering the issue and notified Irregular as well as the affected organisations. The company is also working with independent evaluator METR on a third-party review and plans to improve oversight, reinforce evaluation infrastructure and introduce additional safeguards for future cybersecurity testing. It added that the versions used during the exercises did not include the monitoring systems and safety classifiers deployed with publicly available Claude models.
Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.