1 comment

[ 5.9 ms ] story [ 14.9 ms ] thread
Anthropic was testing Claude in controlled cybersecurity exercises designed to see whether the models could identify and exploit vulnerabilities.

However, some testing environments were accidentally connected to the internet. As a result, Claude gained access to systems belonging to real organizations.

The important detail is that Claude wasn't intentionally sent to attack these companies. The models apparently believed the systems they encountered were part of the testing environment.