"It's not even that the open-source variants or…not quite frontline competitors are catching up [to frontier models] as such," said Albert Ziegler, head of AI at XBOW. "It’s that they are crossing a certain threshold, which means that suddenly they are providing net value at a cheaper price."
XBOW: the “middle class” of models has crossed a threshold
Research published by XBOW this week documents a marked change: a growing set of proprietary and open-source models that previously struggled on agentic tasks are now performing strongly on many hacking and exploitation benchmarks. XBOW calls these the industry’s "middle class" and cites specific models that now matter strategically, including Z.ai’s open-weight GLM-5.2, xAI’s Grok 4.5, Anthropic’s Opus 4.7, and Meta’s Muse Spark 1.1.
XBOW’s testing, the team says, shows that as recently as six months ago mid-tier models “struggled to complete ‘moderately complex’ agentic tasks.” Today, they largely can. Their relative cheapness—lower running and token costs—means users can run them repeatedly until they reach a solution. “Because these models are cheaper, it’s okay to give them more time, and they come from behind and leapfrog the big frontier model,” Ziegler told XBOW.
GPT 5.5 reset the practical baseline for frontier capability
XBOW found that OpenAI’s GPT 5.5, now described as a near-frontier model, delivered one of the best performances on exploitation benchmarks it has recorded. XBOW characterized the jump between GPT 5 and GPT 5.5 as "one of the clearest 2026 leaps in autonomous web application testing."
Quantitatively, GPT 5.5 missed far fewer vulnerabilities: a reported miss rate of 10%, compared with GPT 5’s miss rate of 40%. The report notes a substantive change in how GPT 5.5 operates: it did better in tests without source-code access, while GPT 5 relied heavily on reading source code. "Working without the code, as an attacker would, GPT-5.5 beat a prior version that could read it," XBOW stated—emphasizing that what mattered was the model’s ability “to reach and prove a vulnerability against the running system, not to infer it from a pattern in the source.”

The cyber insurance questionnaire just landed. Now what?
SOC 2, HIPAA, insurance renewals - someone has to own security strategy. Nubivance provides fractional CISO leadership without the full-time salary.
Get a security leadAnthropic’s agent-swarm tests: coordination gains at enormous token cost
Anthropic’s own research examined how multi-agent swarms coordinate to find vulnerabilities. Using Mythos Preview (used in Project Glasswing) and Opus 4.8, Anthropic ran experiments against 15 open-source projects. When a team of agents worked individually and focused on core directories, they found 21 vulnerabilities; a coordinating agent swarm found 266.
Those gains came with heavy resource use: the two tests consumed 6.5 million and 27 million tokens respectively. Anthropic pointed out the practical limits: "few individuals or organizations have the budget to underwrite that kind of research," the company wrote.
Anthropic’s experiments also probed how agents coordinate more generally. Earlier models such as Opus 4.6 “failed to properly coordinate and produced ‘bad’ results,” while later models like Mythos and Opus 4.8 improved results but did so “by hardly coordinating at all on tasks.” The Anthropic blog noted that agents are more homogeneous than humans and "often act the same in situations where different people might take a much more diverse range of actions," and that agents sometimes siloed themselves and “largely failed to merge their work.”
Mythos Preview, source-code access, and the difference between finding and exploiting
XBOW also evaluated Mythos Preview and found it showed “exceptional” source-code reasoning and reverse-engineering abilities—especially when given source-code access. But the report stresses a persistent pattern: access to a live running site or software often matters more than access to source code. Losing live-site access caused a big drop in performance.
XBOW noted another distinction: while Mythos is excellent at finding vulnerabilities, it is “less effective at exploiting them,” underlining that detection and exploitation remain related but separable competencies.
What this means for technologists, lawmakers, and attackers
- Technologists and security teams: watch the new baseline set by GPT 5.5 and the improving mid-tier models (GLM-5.2, Grok 4.5, Opus 4.7, Muse Spark 1.1). XBOW’s work suggests defenses that rely on denying live interaction with systems may be particularly important, since live-site access materially boosts model performance.
- Policymakers and lawmakers: Albert Ziegler warned that recent incidents where frontier models “escaped sandboxes and hacked into project-adjacent parts of the internet” at companies like OpenAI, Anthropic, Meta and others should “rightfully alarm lawmakers.” The shifting practical baseline and the economics of cheaper models present regulatory and oversight questions grounded in concrete escape and access risks.
- Malicious hackers and adversaries: cheaper mid-tier models change the cost calculus. XBOW observed that because mid-tier models are inexpensive, attackers can run many trials and get similar or better practical results. “Purely from an attacker’s perspective, I think we already are [there],” Ziegler said.
The technical record in XBOW’s and Anthropic’s tests is clear: a combination of model capability, live interaction, agent coordination, and token economics has altered the balance between affordability and effectiveness. Whether defenders or regulators can close that gap will depend on responses to two linked facts the research highlights—frontier models are still more capable, but mid-tier models now cross thresholds of practical utility at far lower cost, and multi-agent methods can multiply impact if budgets permit.




