Skip to main content

Tag: ai guardrails

2 articles

Researcher in lab setting with AI equipment and tools.

Researchers Expose Weaknesses in AI Guardrails Against Cyberattacks

Researchers found that AI guardrails against cyberattacks are surprisingly easy to bypass, with attackers often simply telling the model they're allowed to perform a certain action - and it complies. Simple tactics like reframing requests and claiming certain roles reliably trick AIs into assisting with malicious activities.

Analyst 207
Claude Fable 5 model interface on a laptop in a clean room setting with research instruments.

Anthropic Unveils Safer AI Model Fable 5

Anthropic has just unveiled Claude Fable 5, a cutting-edge AI model that's designed with safety in mind, building on the same powerful technology as its predecessor Mythos but with robust guardrails to prevent misuse. This latest release aims to tip the scales in favor of those who can harness its potential responsibly.

Analyst 207