Microsoft has unveiled its first in-house cybersecurity model, MAI-Cyber-1-Flash, trained to hunt for flaws in code and immediately suggest a fix. The company says it beats offerings from Google, OpenAI and Anthropic on the CyberGym benchmark — at roughly half the cost per task.
Alongside the model, Microsoft launched Perception, a platform built around three AI agents: Red, Blue and Green. Red simulates an attack and maps out where a real hacker would strike. Blue sorts through the findings and ranks which flaws actually matter. Green writes the patch and ships the fix — no human required at that last step.
Microsoft AI CEO Mustafa Suleyman said the model "beats Gemini, GPT and Mythos on the industry's primary benchmark." Security engineering lead Dave Weston said tasks that once took hours now take minutes, with MAI-Cyber-1-Flash already handling around 90% of routine security checks on its own.
Microsoft is entering a market where rivals already have a head start: Anthropic launched Mythos in April, and OpenAI shipped Daybreak in May. Microsoft's edge is distribution — the tools plug straight into Azure and its existing enterprise stack, reaching customers who are already paying for its cloud.
There's a catch worth watching. The more autonomy these defensive agents get, the bigger the question of who's accountable if a Green agent ships a bad patch on its own. Microsoft hasn't yet spelled out its safeguards for that scenario in production.



