
Microsoft Launches MAI-Cyber-1-Flash, a Security Model That Beats Mythos and GPT on CyberGym
MAI-Cyber-1-Flash runs inside MDASH and ships into production immediately. Perception adds three agentic security teams. Both are Microsoft firsts, announced in San Francisco on July 27.
MAI-Cyber-1-Flash launched July 27. Microsoft's first dedicated cybersecurity model runs inside MDASH, the company's multi-agent vulnerability identification and remediation harness, and ships into production immediately. At the same event in San Francisco, Microsoft also announced Perception, an agentic security platform that deploys teams of AI agents to monitor, triage, and patch threat vectors continuously.
MAI-Cyber-1-Flash Scores 96% on CyberGym, 12 Points Ahead of Mythos
CyberGym is the gold standard. Microsoft's combined system, MDASH with MAI-Cyber-1-Flash plus GPT-5.4, scores 96% on the benchmark, outscoring Mythos, Gemini, GPT-5.5 Cyber, and GPT-5.6 Sol, which all land between 83% and 86%. MDASH is the same multi-agent harness that already scans every Windows commit for vulnerabilities, now carrying a model purpose-built for security tasks. MAI-Cyber-1-Flash handles roughly 90% of those tasks; GPT-5.4 takes the hardest 10%. The split delivers a 50% cost reduction against Microsoft's previous MDASH stack of GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex.
Perception Assigns Three Agentic Teams to Every Security Workflow
Perception deploys red, blue, and green agent teams. Red teams simulate attacks and surface the vulnerabilities a threat actor would pursue. Blue teams detect and triage existing bugs. Green teams take corrective action. Microsoft's security lead for Perception, Dave Weston, described the platform as compressing what once took hours of manual work across application security hunters and remediation engineers into minutes, producing a discovery, a detection, a posture fix, and a code fix in a single pass. Perception opens to enterprise preview on November 3.
Anthropic Mythos and OpenAI Daybreak both entered the market before Perception. Mythos reached select organizations through Anthropic's Glasswing program earlier this year; OpenAI launched Daybreak in May. Microsoft's answer is a purpose-built security model with benchmark data, one that ships inside a harness covering discovery, detection, posture fixing, and code remediation in a single workflow. Microsoft argues that owning the model, the training data, and the harness gives it an advantage that companies licensing third-party models cannot match.
No third-party model gets Microsoft's data advantage. MAI-Cyber-1-Flash is a compact security model derived from the MAI-Thinking-1 lineage; Microsoft built it on what it describes as an unmatched dataset, 100 trillion daily signals across identity, endpoint, cloud, and network, plus a full record of real exploits and remediations from 1.6 million enterprise customers. AI finding security vulnerabilities at scale is still a developing discipline, and whether purpose-built models like MAI-Cyber-1-Flash outperform general-purpose frontier models on real-world tasks rather than benchmarks is a question the industry has not resolved.