Skip to main content
CybersecurityHacking

Anthropic's Opus 5 Bolsters Defenses Against Prompt Injection Attacks

Laboratory workbench with computer equipment and papers, focusing on an empty laptop screen.

Anthropic’s Opus 5 reduced the chance of a successful prompt‑injection exploit on the IPI benchmark to 2.0% within 15 attempts, down from 5.5% for Opus 4.8 — and just 0.2% on a single attempt compared with 0.5% for the prior release.

Opus 5’s gains against Opus 4.8 on the IPI benchmark

The most direct comparison in the published results is between Opus 5 and Opus 4.8 on the IPI benchmark. Opus 5 cut the measured success rate for an attacker from 5.5% to 2.0% when given up to 15 tries. On a single attempt the improvement was smaller but still measurable: 0.5% down to 0.2%. Those numbers are presented as probabilities of an attacker succeeding under the benchmark’s testing conditions.

Opus 5 relative to Sonnet 5 and Mythos 5

Opus 5 also improved on two other named models in the set: Sonnet 5 and Mythos 5. Sonnet 5 recorded 5.9% success at k=15, and Mythos 5 recorded 2.6% at the same threshold. Against that field, the results show Opus 5 as the most robust model evaluated on the benchmark.

Opus 5 versus non‑Claude models, led by Muse Spark

On the same benchmark Opus 5 outperformed all non‑Claude models included in the comparison. The most robust non‑Claude model cited was Muse Spark, which achieved 16.5% probability of a successful attack within 15 attempts — more than eight times the rate measured for Opus 5 (2.0% at k=15). That gap frames Opus 5 not merely as a modest improvement but as markedly more resistant in these tests than several competitors.

Where GPT 5.6 and its variants sit: Sol, Terra, Luna, and GPT 5.5

The GPT 5.6 family showed a range of robustness in the same measurements. The Sol variant of GPT 5.6 was comparable to the earlier GPT 5.5 on the IPI benchmark, with Sol at 20.0% versus GPT 5.5 at 20.8% within 15 attempts. Those figures make Sol roughly ten times as likely to be successfully attacked as Claude Opus 5 at 2.0% (20.0% versus 2.0% at k=15). Other GPT 5.6 variants were less robust: Terra at 30.4% and Luna at 43.9% within 15 attempts. On single‑attempt performance, a single attempt against GPT 5.6 Sol succeeded 3.1% of the time — higher than the 2.0% chance an attacker achieved against Opus 5 after fifteen attempts.

What this means for technologists, procurement leaders, and adversaries

  • Technologists and security teams: The benchmarked numbers give concrete, comparable failure rates to weigh when selecting or hardening models. Opus 5’s 2.0% rate at k=15 — and 0.2% on a single attempt — is the best reported in this set and may influence model choice when resisting prompt injection is a priority.
  • Enterprises and procurement leaders: The divergence between Opus 5 (2.0% at k=15) and other offerings such as Muse Spark (16.5%) or GPT 5.6 Sol (20.0%) highlights performance differences that can affect risk calculations and contractual expectations for robustness against adversarial prompts.
  • Adversaries and threat actors: The numbers signal where effort might pay off: variants with higher measured success rates — Terra at 30.4% and Luna at 43.9% within 15 attempts — remain more amenable to successful prompt‑injection under the benchmark’s test conditions than the most robust models in the set.

The report closes with a candid observation about limits and progress: "We know that preventing prompt injection is impossible in the general case. But we are getting much better at blocking it in specific cases." The published benchmark data supports that claim by showing measurable reductions in success rates for some models — most notably Anthropic’s Opus 5 — while also documenting substantial variance across vendors and model families.

Original post