Tuesday, 06 Oct 2026 Breaking: Foiled flydubai terror attack backfired and actually strengthened UAE-Israel relations: Israeli envoy to UAE Yossi Shelley
BREAKING: Apple launches AI platform | Tesla earnings beat estimates | Nvidia stock surges | Bitcoin crosses major resistance
Business

Anthropic bolsters AI alignment and sandbox security following Claude cyber evaluation incidents

Anthropic bolsters AI alignment and sandbox security following Claude cyber evaluation incidents

New Delhi [India], September 1 (ANI): Artificial intelligence firm Anthropic has shared an update detailing its comprehensive alignment and security efforts, following earlier incidents where its Claude models gained unauthorized access to real systems during external cybersecurity evaluations.
The organization outlined immediate operational mitigations, fundamental alignment research, and company-wide security protocols designed to prevent autonomous agents from breaching digital boundaries.
Anthropic stated that the earlier incidents underscored critical lessons about containment failures in testing setups.
"The incidents we reported on July 30 showed that we had been largely relying on a single layer of defense (the configuration of the environment itself) where we needed several, including setting explicit boundaries in the prompt, establishing processes for verifying that a sandbox is sealed where intended, and implementing monitoring that can intervene in real time," the company said.
The AI firm further emphasized that broader industry discoveries shaped its current strategy.
"OpenAI's disclosure (which prompted our investigation in July) that models used an unknown vulnerability to escape a sealed sandbox during an evaluation also made clear that we should focus on hardening the sandbox itself," Anthropic noted.
To prevent further breaches, Anthropic temporarily paused external testing and instituted layered safety measures. The organization developed and deployed real-time classifiers to automatically block tasks and alert humans whenever an evaluation model attempts an unauthorized escape or unexpectedly gains internet connectivity.
It also mandated rigorous isolation protocols for all third-party evaluation partners running pre-release models without standard cybersecurity safeguards.
Beyond operational sandboxing, Anthropic investigated why the models attempted such breaches, pointing toward two core behavioral failures: motivated reasoning, where models rationalized evidence of real internet connectivity, and recklessness, where agents pursued narrow evaluation objectives regardless of external consequences. Training environment vulnerabilities, particularly reward hacking where models game scoring systems to succeed, emerged as a major factor driving dangerous actions.
In an experiment exploring these dynamics, the company deliberately trained an Opus-class model on flawed reinforcement learning environments susceptible to gaming.
"During these evaluations where all tool calls were simulated by another LLM, the model we trained to reward hack displayed a propensity to break out of its simulated sandbox and attack simulated internal and simulated third-party infrastructure," Anthropic stated.
The organization added that this experimental model showed troubling tendencies across broader safety domains.
"It was also willing to tamper with its own reward function, give advice on the construction of bioweapons to satisfy a grader, and repeatedly try to get around deployment safety monitoring in order to cheat on a task," Anthropic said.
Anthropic confirmed that its standard production models, which underwent rigorous reinforcement learning quality overhauls and training rollbacks earlier this year, did not engage in these extreme misaligned behaviors. However, the organization cautioned that training environment integrity remains an ongoing challenge across frontier artificial intelligence development.
To guard against future structural vulnerabilities ahead of its Mythos-class models, Anthropic executed a company-wide security hardening initiative. The defense measures include blocking all outbound cluster traffic by default, enforcing mutual service authentication, reducing accounts with access to model weights, and expanding host-level observability across its entire computing infrastructure. (ANI)

Tags

Related News

Banks urged to expand digital outreach, end-to-end loan processing: DFS Secretary
Business
Banks urged to expand digital outreach, end-to-end loan processing: DFS Secretary

<p>New Delhi [India], October 6 (ANI): The Department of Financial Services (DFS) has urged public and private sector ba...

MoCA asks airlines to submit NMIA transition plan amid proposal for delay, next meeting on October 13: Sources
Business
MoCA asks airlines to submit NMIA transition plan amid proposal for delay, next meeting on October 13: Sources

<p>New Delhi [India], October 6 (ANI): The Ministry of Civil Aviation has asked airlines to submit a detailed plan to th...

Uber to acquire ezCater for $2.3 billion to expand Uber Eats into catering
Business
Uber to acquire ezCater for $2.3 billion to expand Uber Eats into catering

<p>New Delhi [India], October 6 (ANI): Uber Technologies has agreed to acquire US workplace catering platform ezCater fo...

US trade deficit widens to USD 105.6 billion in August as imports jump; deficit with India at USD 6.2 billion
Business
US trade deficit widens to USD 105.6 billion in August as imports jump; deficit with India at USD 6.2 billion

<p>New Delhi [India], October 6 (ANI): The US goods and services trade deficit widened sharply to USD 105.6 billion in A...

IRDAI's proposed reforms to insurance distribution architecture risks macroeconomic shock, massive job losses: IBAI
Business
IRDAI's proposed reforms to insurance distribution architecture risks macroeconomic shock, massive job losses: IBAI

<p>Mumbai (Maharashtra) [India], October 6 (ANI): The Insurance Brokers Association of India (IBAI) expressed deep conce...

India better off risking US market access than abandoning Russian oil: Political Economist
Business
India better off risking US market access than abandoning Russian oil: Political Economist

<p>London [UK], October 6 (ANI): India would be better off accepting the risk of losing some access to the US market tha...