Skip to main content

Tag: llms

1 article

Laboratory workstations with computers and notes surround a large monitor displaying a complex neural network diagram.

LLMs' Safety Defense Found Thin and Vulnerable

Researchers made a startling discovery on Qwen3-4B, finding that a mere 50 neurons - just 0.014% of the model's feed-forward neurons - control its safety defense, and removing them dramatically changed the model's response to harmful prompts. Disabling these neurons altered the model's refusal format in 80% of 520 standard harmful-prompt benchmarks.

Analyst 207