Anthropic says Claude accessed real systems during cyber evals
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing.…
RT @dok2001: Twice in nine days. OpenAI's models chained a zero-day to get out of an eval environment. Anthropic just found three incidents…
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…
Read full post
We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v
RT @AnthropicAI: In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from w…
These companies really are just different-themed versions of much the same thing. Anthropic gives us a cute little graphic when their AI breaks out of its sandbox and hacks other companies.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…
Read full post
We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v
RT @Miles_Brundage: Our innocent harness misconfig, their egregious misalignment
This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing!
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…
Read full post
We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v
RT @Sauers_:
OpenAI had a damning security incident where their under development AI escaped the sandbox environment and attempted to hack another company (HuggingFace) For some weird reason Anthropic decided to share a similar incident from 3 months ago, only NOW. Something smells off…
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…
Read full post
We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v
T-minus 5 days until chinese model labs will have their own cost efficient model breakout
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…
Read full post
We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v
RT @AnthropicAI: In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from w…
RT @tszzl: both of the leading labs have had serious loss of control incidents. there will be serious coping about this from both sides and…
RT @Sauers_: - you're Claude - "hack this fictional company" - can't figure out how to hack the simulation. let me try the internet. - "htt…
RT @Sauers_:
RT @Laz4rz: or "they dont know my creators have to now catch up in fearmongering too" now
AI getting over the internet and gaining unauthorized access! Looks like a plot of a Hollywood movie
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing.…
Read full post
We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. https://t.co/dKFCdpKd9v
RT @kpolley: “AI Meltdown” is the scenario where the agent goes off the rails and forgets/ignores all previous instructions and soft guardr…
Original title: Investigating three real-world incidents in our cybersecurity evaluations
Samuel Times preserves the original link so every selection remains auditable.


