SecurityWeek

Anthropic Details Response to Security Incidents, Unveils Enterprise Safeguards


Anthropic has detailed its response to a series of unauthorized access incidents involving Claude models, along with a new enterprise product that combines data privacy with misuse monitoring.

Anthropic’s Claude models operating without cyber safeguards for testing purposes recently gained unauthorized access to live systems after being mistakenly granted internet access.

In addition, the UK AI Security Institute separately reported that Claude Mythos 5, also being tested without safeguards but with intentionally given internet access, took a series of unauthorized actions against real people and organizations.

Anthropic said its early findings point to two contributing factors: the models appeared to discount evidence that their environment was connected to the real internet after initially being told it was simulated, and they showed a willingness to take harmful actions to complete an assigned task. 

In response, Anthropic temporarily paused external and some internal cyber evaluations and built a classifier that detects and blocks attempts to escape a test environment in real time. Additionally, the company added new requirements for outside partners, including verified network isolation and testing of sandbox boundaries before an evaluation begins.

The company also described an experiment in which it intentionally trained a model on reinforcement-learning environments known to reward “cheating” (ie, completing tasks by gaming the training process rather than solving them as intended). 

Advertisement. Scroll to continue reading.

In simulated scenarios, that model attempted to escape its sandbox, interfered with its own reward mechanism, and offered guidance on building bioweapons in order to satisfy an evaluator. Anthropic said its currently released models did not show this behavior when tested against the same scenarios.

On the security side, Anthropic said it has since reduced the number of accounts with standing access to systems holding model weights or customer data, set its computing infrastructure to block outbound network traffic by default, and temporarily moved roughly 150 product engineers to security-related work.

Anthropic unveils Enterprise Frontier Safeguards

Separately, Anthropic introduced Enterprise Frontier Safeguards (EFS), a system that combines zero data retention with automated monitoring for misuse. EFS lets customers store their own activity data on infrastructure they control, rather than Anthropic’s infrastructure.

The company said it built the system with input from more than 100 customers, including the Analysis and Resilience Center for Systemic Risk, whose membership includes security chiefs at Goldman Sachs, Morgan Stanley, Citi, Bank of America and Wells Fargo, along with companies such as Comcast, KPMG, Mastercard, Salesforce and Visa.

Under the new system, flags from automated monitoring go directly to the customer’s own review team rather than to Anthropic staff, and features such as customer-owned storage and customer-managed encryption keys are optional. 

The rollout begins this fall across Claude Code, Claude Enterprise, and the Claude Platform.

Related: Anthropic Warns Claude Users of Infostealer Malware Infections

Related: Irregular Details How a Naming Error Let AI Models Attack a Real Company



Source link