Friday, 31 July 2026Independent intelligence from the feedEditor
Samuel TimesSignal over noise
Front Page/AI
AI cybersecurity evaluations · Signal 9.0

Anthropic says Claude accessed real systems during cyber evals

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing.…

Hand with padlock and key on detailed security graphic
Image from the primary source

RT @dok2001: Twice in nine days. OpenAI's models chained a zero-day to get out of an eval environment. Anthropic just found three incidents…

Media posted by @harpaljadeja

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…

Read full post

We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v

These companies really are just different-themed versions of much the same thing. Anthropic gives us a cute little graphic when their AI breaks out of its sandbox and hacks other companies.

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…

Read full post

We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing!

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…

Read full post

We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v

OpenAI had a damning security incident where their under development AI escaped the sandbox environment and attempted to hack another company (HuggingFace) For some weird reason Anthropic decided to share a similar incident from 3 months ago, only NOW. Something smells off…

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…

Read full post

We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v

T-minus 5 days until chinese model labs will have their own cost efficient model breakout

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…

Read full post

We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v

RT @tszzl: both of the leading labs have had serious loss of control incidents. there will be serious coping about this from both sides and…

RT @Sauers_: - you're Claude - "hack this fictional company" - can't figure out how to hack the simulation. let me try the internet. - "htt…

AI getting over the internet and gaining unauthorized access! Looks like a plot of a Hollywood movie

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…

Read full post

We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v

Continue with the primary sourceOpen www.anthropic.com

Original title: Investigating three real-world incidents in our cybersecurity evaluations

Samuel Times preserves the original link so every selection remains auditable.