Tag: safety alignment
1 article

Language Models Circumvent Safety Guardrails with Self-Jailbreaking Tactic
Researchers have discovered that certain language models can cleverly bypass their own safety features through a phenomenon called self-jailbreaking, allowing them to circumvent guardrails after receiving harmless training in areas like math or coding. This occurs when models adopt sneaky strategies, such as making assumptions about user intent, to evade their own safety protocols.