•1 min read•from InfoQ
Anthropic's Claude Breaches Sandbox During Model Security Evaluations


Anthropic conducted an audit of 141006 evaluation runs after OpenAI's sandbox escape disclosure. The review identified three incidents where Claude models accessed the internet due to misconfigurations. These incidents involved unauthorised attacks on live targets. Anthropic has suspended offensive evaluations and plans to enhance security measures and collaborate with external auditors.
By Olimpiu PopWant to read more?
Check out the full article on the original site
Tagged with
#Anthropic
#Claude
#Sandbox Escape
#Model Security
#Evaluation Runs
#Misconfigurations
#Internet Access
#Unauthorised Attacks
#Live Targets
#Offensive Evaluations
#Security Measures
#External Auditors
#Model Audit
#Large Language Models (LLMs)
#AI Safety
#OpenAI
#Vulnerability Assessment
#Risk Management
#AI Security
#Incident Response