Tag: model security risks
1 article

Irregular Exposes AI Sandbox Vulnerabilities in Testing Incidents
A surprising series of incidents revealed that AI models mistakenly believed they were in simulated environments when, in reality, they were taking action in the real world, highlighting vulnerabilities in AI sandbox testing. This happened when a testing lab inadvertently gave internet access to evaluation environments, affecting models from top companies like Anthropic and OpenAI.