Microsoft says that the Frontier Offensive Research & Generative Exploitation Lab, known by the abbreviation FORGE, has moved from testing AI’s ability to find difficult vulnerabilities to studying what is required to turn these discoveries into shippable fixes. Between May and September 2026, the lab helped discover vulnerabilities in Windows that were assigned 140 CVE identities, 52 of which were addressed in the September 2026 security updates.
The work also extended to open-source software; FORGE teams submitted approximately 155 reports that were internally validated across 23 projects, including the Linux kernel. According to Microsoft, 93 reports across 14 projects or project families had received documented acknowledgment or acceptance from maintainers at the time the material was prepared. One Linux report submitted through the Linux Foundation’s Akrites initiative also became the first report from the initiative to result in a patch being merged into the Linux kernel.
From Maximum Capability to Operation at Scale
The first lesson is that a model’s success in finding a complex flaw does not guarantee that it will produce fixes at a comparable pace. Increasing the number of validators may raise the number of candidates, but it does not necessarily increase the number of confirmed findings or published fixes. If reports arrive faster than the Microsoft Security Response Center can review them, a queue builds up that consumes experts’ time, especially when reports are duplicated or lack reproducible evidence.
For this reason, FORGE focuses on a reproducible result, not solely on the number of scans or reports. The lab uses the multi-model MDASH platform to organize the work, along with exploit or proof-of-concept generators, test harness construction, and the search for inputs that trigger faulty behavior. Microsoft says that an internal project relying on deterministic algorithms, such as abstract syntax tree analysis, reduced duplicate reports by about 45% across multiple scans of the same code.
Spending on Reducing Uncertainty, Not on Increasing Text
Microsoft believes that measuring efficiency solely by the number of tokens produced by the model leads to a misleading metric. A short report may be ambiguous and costly to investigate, while a longer analysis may justify the execution path and reduce the total cost. The practical decision is to determine what the investigator lacks—the function caller, the build configuration, a runnable reproducer, or the causal explanation—and then direct the next step specifically toward closing that gap.
MDASH combines advanced models with distilled models, specialized validators, and code-analysis tools. It can direct routine tasks to less expensive models and escalate unresolved questions to more powerful models. However, Microsoft emphasizes that this policy remains a hypothesis requiring measurement, because early filtering may exclude real vulnerabilities.
Validation and Fixing as a Continuous Learning Loop
According to the proposal, validation and fixing should not be treated as a later phase after vulnerability discovery. Every finding should pass through automated validation, human review, patch development, and regression testing, with the evidence from each stage fed back into the system. Even validation failure can be useful if its cause is recorded, such as an unreachable path, an incorrect build configuration, the absence of attacker control over inputs, or a duplicate report.
In an experiment on the Linux kernel, validation agents produced supporting evidence for 627 findings, while the average cost of creating a proof of concept for 182 confirmed crash findings was $3.61 in model cost and 21.5 minutes per successful case. In six cases testing the possibility of local privilege escalation through automated exploit generation, the average was $8.56 and 25.4 minutes. The evaluations used GPT-5.5 and did not include the costs of initial scanning, failed cases, human investigation, or patch preparation.
What Does This Mean for Security Teams?
The most important reading of these results is that the success metric for agentic vulnerability-discovery systems should not stop at the number of findings. Microsoft suggests tracking the number of confirmed flaws, duplicate and rejected reports, queue age, the time from discovery to fixing, as well as model costs and human review time. Work on open-source projects also requires respecting each project’s disclosure, review, and fixing processes, rather than simply sending more reports.
In practice, the source identifies a shift in the bottleneck: the ability to find a flaw has become only part of the problem, while linking research to build environments, binaries, configurations, test tools, and CI/CD systems is becoming more important. The limits of the results also remain clear; the figures for validation cost exclude important human stages, and turning research results into an accepted fix depends on maintainers and the context of each project. Therefore, the material does not prove that automation replaces engineering review; rather, it presents automation as a way to reduce repetitive work while keeping responsibility for security judgment and fixing with humans.