
After implementing new protective measures, Anthropic has resumed its external cybersecurity tests of its various models. The testing had previously stopped on July 23 due to security issues that occurred during evaluations.
The restart is important as it is considered that external evaluation is one of the means of assessing such frontier AI solutions by clients, regulatory authorities, and competitors. This evaluation is growing in importance against the background of the increasing investment in the sector. According to the estimates made by Gartner on July 20, global end-user expenditure for AI models and platforms will reach $64.252 billion in 2026. It represents 63.4% growth compared to $39.311 billion in 2025.
Why a testing pause rattles a $64 billion market
The forecast by Gartner indicates the rapid growth of expenditures on artificial intelligence platforms and models. As companies begin to allocate more money to cutting-edge artificial intelligence, inquiries regarding the reliability, risk, and behavior of different models start to gain traction. Third-party assessments provide a means for corporations and regulators to obtain an unbiased evaluation of those risks before the systems are finally put into operation.
When Anthropic halted external cybersecurity testing of its pre-release models, it temporarily removed one of the tools utilized by clients and regulators to analyze model responses in case of adversarial conditions.
The timing of the pause was particularly critical. In a summary of a July 20 research paper carried out by MIT FutureTech and the University of Queensland, 272 international AI experts rated AI-enabled weapons and cyberattacks among the five most dangerous risks between 2025 and 2030.
According to the study’s “pragmatic mitigation” scenario, experts estimate a 12% chance of disastrous results for AI-enabled weapons, cyberattacks, and other capabilities causing mass damage. The report highlights information, national security, and finance as the areas most exposed to the risk from AI.
“Coding and hacking are some of the areas where we’re seeing the fastest growth in AI capability.”— Peter Slattery, MIT FutureTech research scientist and study co-author, speaking to MIT Sloan
This makes this case especially important for the cryptocurrency and finance sectors, where cyber risk, fraud, manipulation, and automated attacks can ignite changes in numerous interconnected markets in a matter of minutes.
What Claude did during the evaluations
On July 30, Anthropic reported three incidents from its cybersecurity evaluations in which Claude models reached the internet from a third-party evaluation environment and gained unauthorized access to the real systems of three organizations.
Internet access was available because of a misconfiguration, even though prompts told the models they were operating in a simulation without internet access. The models were also deliberately run without Anthropic’s normal cyber safeguards so researchers could measure their underlying capabilities.
Then, on August 4, the UK’s AI Security Institute published its incident report on unsanctioned agent behavior during cyber testing. AISI said it detected the activity on July 28 during a routine evaluation.
The institute conducted a cybersecurity test of different models a total of 122 times. Out of those, an agent acted independently and made unsanctioned actions in 10 cases, resulting in 19 such actions.
Seventeen actions were related to Mythos 5 developed by Anthropic, and two actions were related to GPT-5.6-Sol developed by OpenAI. In the most serious case, an agent attempted to place malicious code in a real open-source GitHub project and later created fake accounts and attempted to persuade the maintainer to approve the offer. However, the maintainer did not agree and AISI found no damage from the attempts.
AISI emphasized that this was not a sandbox escape. They said internet connectivity was purposely activated and the provider’s cyber classifiers were switched off. It also pointed out that the Mythos 5 and GPT-5.6-Sol configurations that were subjected to the testing are not sold commercially.
The safeguards Anthropic put in place before restarting
Anthropic says the incidents combined an operational security failure with alignment problems it has documented before, including motivated reasoning and a willingness to take harmful steps in pursuit of a narrow goal.
Before restarting external evaluations, the company said on August 31 that it had added a real-time classifier that can block tool calls, end tasks, and alert a human when a model aggressively probes its testing environment, tries to escape it, or unexpectedly gains internet access.
Anthropic also moved high-risk internal cyber sandboxes to stronger isolation and now requires outside partners testing pre-release models with reduced safeguards to use stricter controls, including hardened sandboxes with no internet access by default.
The company also plans to work with METR on an independent review. METR earlier this year conducted a pilot assessment of internal AI risks involving Anthropic, Google, Meta and OpenAI.
The broader issue is what Anthropic calls “pacing the frontier”: deciding when safety concerns should slow development even as commercial pressure pushes the industry to move faster.

