On October 4, 2026, Google temporarily stopped accepting product-flaw reports for its open-source bug bounty while keeping supply-chain submissions open. Four years earlier, it had offered up to $31,337 to attract vulnerability reports from outsiders. One part of that invitation was now on hold.
Key takeaways
- Google launched its OSS Vulnerability Reward Program in 2022 with rewards of up to $31,337.
- Google’s separate broader Vulnerability Reward Program paid $10 million to 632 researchers in 2023.
- On October 4, 2026, Google paused OSS product-vulnerability submissions while leaving supply-chain and outstanding reports unaffected.
- Google said it planned to provide an update on the OSS product-report pause by Q1 2027.
- CWE-Bench-AA contains 120 held-out auditing and patching tasks for evaluating security agents.
In 2022, Google launched its OSS Vulnerability Reward Program to reward reports about open-source projects. Its broader, separate Vulnerability Reward Program paid $10 million to 632 researchers in 2023. The two programs had different scopes, but both used rewards to direct outside attention toward flaws a vendor might otherwise miss.
AI agents can generate plausible bug reports at lower cost without necessarily reducing the work needed to verify them. Bug bounties were built to compete for scarce discovery; their harder constraint is becoming the capacity to authenticate, reproduce, deduplicate and act on what researchers submit.
The bounty bought attention when finding a flaw was expensive
A researcher once had to spend substantial effort inspecting a target, developing a credible finding and preparing it for coordinated disclosure before a vendor would consider a payout. The vendor still had to test the report, but the researcher’s work filtered many weak possibilities before they reached the intake queue. Rewards gave researchers a legitimate place to send a finding, a process for disclosure and a reason to look at the vendor’s software rather than someone else’s.
Google’s 2022 OSS program applied that bargain to open-source projects. A larger reward could make a difficult target worth inspecting, but it could not help a reviewer reproduce a flawed report.
DARPA launched its AI Cyber Challenge in 2023 to develop systems that scan open-source code for security flaws. In April 2026, Bloomberg reported that 360 Digital Security Group had used an AI-powered agent to uncover roughly 1,000 previously unknown vulnerabilities, including flaws in Microsoft Office. Neither effort proves that automated findings are noise. Both show how much easier it is becoming to produce another candidate for review.
Google closed one intake lane, not the program
On October 4, Google said it was temporarily no longer accepting OSS VRP product-vulnerability submissions. Supply-chain reports and outstanding reports were unaffected, and Google planned an update by Q1 2027. Google restricted a particular intake category rather than abandoning outside disclosure or declaring AI-generated research worthless.
The reported reason was an influx of invalid AI-driven reports. Claims that engineers and open-source maintainers were overwhelmed by thousands of sloppy submissions remain unverified. Google’s pause establishes neither its submission volume nor an industry-wide invalid-report rate.
Other operators have encountered the same kind of capacity problem. The curl project planned to end its HackerOne bounty program, citing low-quality AI-generated reports. The Financial Times reported that bug-bounty operators were adding background checks and building AI triage agents. A maintainer receiving a plausible claim must still establish the affected version, reproduce the behavior, assess impact and determine whether someone already reported it. Even a report that fails those checks consumes time before it can be dismissed.
A finding becomes valuable when another person can test it
A prose report can name a real function and describe a plausible failure without establishing that the software behaves as claimed. A stronger submission gives the recipient an affected version, a path to reproduce the behavior, evidence of impact and enough provenance to compare it with earlier reports. Those materials let a maintainer challenge the claim without first reconstructing the researcher’s entire investigation.
OpenAI positioned Codex Security as an agent that finds, validates and proposes fixes for vulnerabilities. OpenAI also said its Aardvark private beta improved signal quality, severity accuracy and false-positive rates. The same tools that make candidates cheap may help test them. OpenAI’s claims do not establish how much reviewer work disappears in another organization’s codebase, but they place validation inside the product rather than leaving it to whoever opens the inbox.
Artificial Analysis and its partners made a related choice in measurement. Their Cyber Index Alliance evaluates agents on finding and fixing vulnerabilities; CWE-Bench-AA includes 120 held-out auditing and patching tasks. As with AI coding benchmarks more broadly, a score is useful only to the extent that the test captures the work a recipient must trust. Finding a suspicious line is different from reproducing a bug, and producing a patch is different from showing that the patch does not break the software.
Someone must still decide what level of impact warrants an emergency fix, who owns the affected project and when disclosure is responsible. Those decisions need a record that another maintainer can inspect.
Faster discovery makes the handoff to repair more consequential
OpenAI’s Daybreak initiative combines models, Codex Security and security partners to help organizations find, validate and patch vulnerabilities continuously. That design connects tasks that a bounty program can otherwise separate into different queues: a researcher submits a claim, a reviewer verifies it, an owner receives it and a maintainer prepares a fix. When any handoff loses the reproduction steps or the reason for a severity decision, the next person has to repeat the work.
A verifier can record the environment and observed behavior; a reviewer can attach the duplicate and severity decisions; a maintainer can test a proposed patch against the same case. That chain is the practical work of operational AI assurance, though current agents have not been shown to complete every step reliably.
Delay also has an adversarial cost. Google’s Threat Intelligence Group reported what it described as the first known case of hackers using AI to discover and weaponize a zero-day. That report does not tell a bounty operator to accept weak submissions faster. It makes accurate, reproducible submissions more useful: a vendor that cannot distinguish a real flaw from a persuasive false alarm spends time on both while an attacker needs only the real one.
Trusted access saves review time by limiting who gets in
Participant screening predates the AI-report surge. In 2020, election-technology company ES&S partnered with Synack to let Synack-vetted professionals test some of its products. Vetting gave ES&S a way to permit outside scrutiny while deciding who could conduct it. The company accepted narrower participation in exchange for a relationship it could govern.
GitHub’s planned two-tier bounty program makes that trade in compensation: lower rewards for public submissions and higher payouts for invite-only researchers amid AI-generated report volume. Apple introduced a submission cap and a 30-day cool-off period, while allowing researchers to request higher quotas. Each rule limits the work entering a reviewer’s queue before the reviewer knows whether a particular claim is valid.
An unknown researcher can find a flaw a vetted group misses, and an open channel can reach parts of a codebase that an invitation list does not. Background checks, quotas and reputation can protect maintainers’ time while excluding useful work. Machine-assisted validation offers another way to keep the door wider, provided it supplies evidence that a human can inspect rather than another confident verdict to investigate.
Frequently asked questions
When will Google say whether OSS product-vulnerability submissions will reopen?
Google said it planned an update by Q1 2027. The piece does not identify decision criteria, a reopening date, or whether the pause will become permanent.
How large was the reported influx of invalid AI-driven reports?
Google’s pause does not establish a submission total or an invalid-report rate. Claims that maintainers faced thousands of sloppy reports remain unverified in the piece.
Does the $31,337 figure mean every valid OSS report could receive that amount?
No. It was the program’s stated maximum reward when Google launched the OSS program in 2022, not a guaranteed payout for each valid finding.
What is a comparable example of limiting researcher access before reports reach maintainers?
In 2020, ES&S partnered with Synack to allow Synack-vetted security professionals to test some ES&S products. That model narrowed participation in exchange for researchers whose access and identity could be governed.
Milestones in AI-assisted vulnerability review
- 2020-08-06 — ES&S partnered with Synack to allow Synack-vetted professionals to conduct penetration testing on some products.
- 2022 — Google launched its OSS Vulnerability Reward Program, offering rewards of up to $31,337.
- 2023 — DARPA launched its AI Cyber Challenge for systems that scan open-source code for security flaws.
- 2026-04 — Bloomberg reported that 360 Digital Security Group had used an AI-powered agent to uncover roughly 1,000 previously unknown vulnerabilities.
- 2026-09-29 — Artificial Analysis launched the Cyber Index Alliance; its CWE-Bench-AA includes 120 held-out auditing and patching tasks.
- 2026-10-04 — Google paused OSS product-vulnerability submissions and said it would provide an update by Q1 2027.
Google’s $31,337 ceiling advertised the value of finding a flaw. Its closed product-flaw lane showed how much a claim still has to prove before that prize means anything.