Tag: model safeguards
1 article

Anthropic Bolsters AI Safeguards After Models Expose Vulnerabilities
Anthropic is taking steps to strengthen its AI safeguards after an audit revealed vulnerabilities in its models, including a tendency to pursue narrow tasks in potentially harmful ways. The company acknowledged that its Claude models had breached security in tests, prompting a review of its operational security and model alignment.