Anthropic says some Claude models gained access to the live internet and hacked three organisations during what was supposed to be an isolated cybersecurity evaluation. The models were tasked with breaking into test machines, but because of a misconfiguration, they reportedly reached real external systems and hacked three unnamed organisations. Anthropic says it reported the incidents and is urging other AI labs to review their own evals for similar failures.
Click here for the official article/release
Disclaimer
The Legal Wire takes all necessary precautions to ensure that the materials, information, and documents on its website, including but not limited to articles, newsletters, reports, and blogs (“Materials”), are accurate and complete. Nevertheless, these Materials are intended solely for general informational purposes and do not constitute legal advice. They may not necessarily reflect the current laws or regulations. The Materials should not be interpreted as legal advice on any specific matter. Furthermore, the content and interpretation of the Materials and the laws discussed within are subject to change.
