Skip to main content
CybersecurityVulnerability Management

Microsoft Unveils AI-Powered Security Model to Outperform Rivals

Sleek computer terminal with futuristic security props on minimalist background.

"This is really quite a remarkable result," Mustafa Suleyman, CEO of Microsoft AI, said after Redmond published benchmark numbers showing a new combination of specialized models and an agentic harness achieved a 95.95 percent success rate on a vulnerability task.

MAI-Cyber-1-Flash, MAI-Thinking-1 and the MDASH harness

On Monday at Redmond’s security event, Microsoft unveiled MAI-Cyber-1-Flash, its first security-specialized model, and said it had been packed inside a bug-hunting harness called MDASH. The company described MAI-Cyber-1-Flash as built on the internally developed MAI-Thinking-1 reasoning model and combined it with a larger system boost from GPT-5.4.

Microsoft executives explained that, inside MDASH, MAI-Cyber-1-Flash handles up to 90 percent of incoming queries — detecting vulnerabilities, producing patches, and confirming that the fixes worked — and then hands the remaining roughly 10 percent of tasks to the larger GPT-5.4 model. Suleyman described GPT-5.4 as "about 10X larger" and said the handoff between models yields both higher measured performance and lower operational cost.

CyberGym benchmark: numbers that drive the narrative

Microsoft cited third-party benchmarking by CyberGym showing the MDASH configuration — MAI-Cyber-1-Flash combined with GPT-5.4, both inside MDASH — achieved a 95.95 percent success rate against real-world vulnerability tasks. Microsoft contrasted that figure with several competitors’ results from the same benchmark: OpenAI’s GPT-5.5 Cyber at 85.6 percent, OpenAI’s GPT-5.6 Sol at 83.6 percent, Anthropic’s Mythos 5 at 83.8 percent on real-world vulnerabilities, and Google’s Gemini 3.5 Flash Cyber in CodeMender at 83.2 percent.

Microsoft also claimed the MDASH approach costs about half the price of other leading commercial models, an efficiency the company attributed to the division of labor between its specialized and larger models.

Project Perception: red, blue and green agents

Alongside the model announcement, Microsoft introduced Project Perception, a system that coordinates three agent types. According to Microsoft, red team agents find and simulate attack paths, blue team agents investigate and determine risk, and green team agents remediate issues. "We need to make sure that the defenders can defend at the scale and the speed of the attackers," Hayete Gallot, executive vice president of Microsoft Security, said when describing the new cyber stack that Perception is meant to deliver.

Perception is positioned as an agentic orchestration layer intended to automate the sequence from discovery and validation through remediation, by allocating tasks across purpose-built agents.

Microsoft Security FORGE Labs and the External Red Team Alliance (EXTRA)

Microsoft also launched Microsoft Security FORGE (Frontier Offensive Research and Generative Exploration) Labs, an AI security research arm led by Microsoft VP of Security Research Taesoo Kim. In parallel, the company announced the External Red Team Alliance (EXTRA), described as a two-part initiative to expand AI safety research through funding and an extended red-teaming network.

Ram Shankar Siva Kumar, Microsoft’s data cowboy and AI red team lead, said in a blog that Redmond’s own AI red team has provided "unrestricted gifts" to 18 university labs across six continents to support AI safety research. "The funding is unrestricted because the objective is not to direct research outcomes toward product requirements or predefined deliverables," he wrote, adding that some universities are studying how models can be attacked or abused in operational environments, while others are exploring how AI systems can assist defenders.

The second component of EXTRA will build a distributed network of specialists — researchers, practitioners and regional experts — to participate in red teaming across specific attack classes, languages, cultural contexts or technical domains that internal teams may not fully cover.

What this means for security teams, procurement leaders, and university researchers

  • Security teams: The MDASH architecture promises an automated flow from detection to confirmed remediation; teams will need to evaluate whether the claimed 90/10 split between specialized and larger models matches their operational realities and whether MDASH integrates with existing workflows.
  • Procurement leaders and CIOs: Microsoft’s claim of roughly half the cost of rival commercial models frames a procurement conversation around total cost of ownership for mixed-model harnesses versus single-model deployments; purchasers will likely seek audited benchmark data and pricing details tied to real-world usage.
  • University researchers and external red teamers: Microsoft’s "unrestricted gifts" to 18 university labs and the EXTRA network create new funding and collaboration pathways for work on both offensive and defensive AI-security problems, with explicit emphasis on domains internal teams may not fully cover.

Microsoft’s announcements bundle product, benchmark and research investments into a single narrative: specialized models plus agentic orchestration can raise measured performance while reducing cost. The company has put names and numbers behind that narrative — MAI-Cyber-1-Flash, MAI-Thinking-1, MDASH, Project Perception, FORGE Labs, EXTRA, and the 95.95 percent CyberGym result — and framed the effort as a new "cyber stack" to let defenders operate "at the scale and the speed of the attackers," in Gallot’s phrasing. How broadly those claims translate into deployed advantage will depend on adoption, independent verification, and the operational tests that follow.

Original story