Cybersecurity analyst monitoring network threat data on screens

How Does a Generator Work in Cybersecurity

October 7, 2026 · 9 min read · By Rafael

Key Takeaways:

  • The National Vulnerability Database logged 60,475 flaws by September 2026, versus 48,185 for all of 2025, a pace driven by AI-assisted discovery.
  • PROMPTFLUX malware, found by Google’s Threat Intelligence Group, uses Gemini to rewrite its own code every hour, defeating signature-based detection.
  • OpenAI’s Daybreak Cyber Partner Program now lets 30 vendors embed GPT-5.5 in customer-facing defenses, and Cloudflare pairs GPT-5.6-Cyber with edge telemetry to find and block flaws.
  • Standalone static analysis carries roughly a 35.7 percent precision rate; pairing LLM reasoning with SAST cuts false positives by up to 91 percent.
  • OpenAI committed $1 billion in subsidized Daybreak access, but defenders note the same week shipped Astra, a model that crossed the “Critical” cybersecurity threshold.

Self-Rewriting Malware and the End of Signatures

The clearest sign that static defenses are losing came from Google’s Threat Intelligence Group, which identified a malware strain called PROMPTFLUX that uses Gemini to rewrite its own code every hour. A signature that matches the binary at 9 a.m. is useless against the version that exists at 10 a.m. This is not theoretical evasion; it is a design choice that makes hash-based and pattern-based detection structurally obsolete against that particular threat.

Cybersecurity analyst monitoring network threat data
Signature-based detection assumes a stable artifact to match. Self-rewriting malware removes that assumption.

Microsoft’s 2026 Digital Defense Report reached a related conclusion: AI has given cyberattackers an immediate advantage, including the first fully autonomous ransomware to compromise real organizations. The report describes the problem as a speed imbalance. Attackers can iterate continuously, while defenders still work through ticket queues, approval chains, and change windows measured in days. The time between a flaw becoming known and that flaw being weaponized has shrunk from weeks to hours.

AI Vulnerability Detection as a Generator

Defenders are responding by using models to generate findings and fixes rather than just scanning passively. OpenAI’s Daybreak Cyber Partner Program, announced in June 2026, allows 30 cybersecurity vendors and service providers to integrate GPT-5.5 into products their customers already use. Vendors described the intended work as vulnerability prioritization, faster incident response, enriched threat intelligence, and automated attack-path analysis, according to BankInfoSecurity’s reporting.

The most concrete deployment is Cloudflare’s Vulnerability Discovery and Remediation service, opened in early access in September 2026. It runs inside Cloudflare Managed Defense and reaches OpenAI’s GPT-5.6-Cyber model through the Daybreak Defense Network. Cloudflare takes a snapshot from its Web Assets inventory and web application firewall to see which routes are live, then a reconnaissance agent maps request paths to sections of the codebase and hunter agents work from that map. Every finding is validated before it receives a risk rating, and that rating rises if production evidence shows heavy traffic or active probing against the affected route.

Two kinds of output come out of that process. Custom firewall rules, scoped to the method, path, and request details needed to reach the vulnerable code, can go up as a stopgap while developers work. The models also draft code patches for engineers to review. Neither a rule nor a patch takes effect without explicit human approval, and Cloudflare states the models cannot apply either on their own. Prompts travel through Cloudflare’s AI Gateway to OpenAI servers, with redaction stripping out context the model does not need, and every proposal must pass checks written outside the model.

The Patch Cycle Cannot Keep Up

The volume problem is now measurable at the vendor level. Microsoft’s September 2026 Patch Tuesday was a record, with the company describing it as close to 1,000 vulnerabilities, according to its CISO guidance post. Independent tallies put the figure at 974 flaws, two of them already exploited in the wild. Microsoft attributed the surge directly to frontier AI models finding more issues in its codebase than human review ever did.

The cost of falling behind is concrete. IBM’s 2026 report puts the global average breach cost at a record USD 4.99 million, with one in four malicious breaches involving AI. Action1’s 2026 Software Vulnerability Ratings Report adds a sharper trend line: disclosed vulnerabilities across the enterprise software categories it analyzed increased 92 percent in 2025 compared with 2024, while critical and high-severity flaws each rose 103 percent, and remote-code-execution vulnerabilities rose 128 percent.

There is a second-order problem hiding inside that volume. In April 2026, NIST reclassified roughly 30,000 vulnerabilities published before March 1 as “Not Scheduled” for enrichment, because CVE volume had outgrown the model NVD was designed around. That means a large backlog of known flaws now lacks the normalized metadata defenders rely on to decide whether a vulnerability applies to their environment. Teams that depend on NVD as their single source of truth are increasingly working from partial intelligence.

Benchmarking AI’s Evasion and False Positives

AI-driven detection faces a major limitation: language models hallucinate. AWS’s Deception Benchmark, made public in September 2026, tests whether AI models can tell real security vulnerabilities from safe code that looks risky, and the early result is that false positives pile up. A model that flags safe code as vulnerable is not free; it burns analyst time and erodes trust in the tool.

The numbers on this trade-off are now documented. Independent research has measured a precision rate of roughly 35.7 percent for standalone static application security testing engines, meaning the majority of flagged findings are false positives. Pairing LLM contextual reasoning with traditional static analysis reduces false positives by up to 91 percent, according to the same research cited by GitHub in its Copilot CLI announcement.

GitHub’s own approach is different: its experimental /security-review command in Copilot CLI replaces the rule engine with LLM inference but restricts output to high-confidence findings only. It targets five high-impact vulnerability classes, injection flaws including SQL injection, cross-site scripting, insecure data handling, path traversal, and weak cryptography, and runs pre-commit inside the terminal, including in air-gapped environments via a bring-your-own-key mode. The trade-off is coverage versus noise. The command does not do CVE matching, taint analysis, or dependency scanning, so it sits alongside CodeQL and Dependabot rather than replacing them.

Astra and the Dual-Use Threshold

The same model that finds a flaw can exploit it. OpenAI’s Astra, designated on September 1, 2026, is the first model to cross the “Critical” cybersecurity capability threshold under the company’s Preparedness Framework, the framework’s highest tier, defined as introducing unprecedented new pathways to severe harm. On ExploitBench, which measures turning a known vulnerability into a working exploit, Astra scored 100 percent against 78.5 percent for GPT-5.6 Sol, per OpenAI’s own benchmark.

That dual-use nature is the core tension of the entire category. A model good enough to find and patch a zero-day in a widely deployed engine is also capable of weaponizing one. OpenAI released Astra with offensive capabilities restricted to secure code review and patching, refusing prompts that request proof-of-concept exploits, with less restrictive settings planned for the vetted Daybreak program. The limitation is real but narrow: the capability remains inside OpenAI, and the restriction is a policy control, not a property of the model.

The timing problem received sharp criticism from practitioners. Neil “Grifter” Wyler, vice president of defensive services at Coalfire, told TechTarget that OpenAI shipped its most powerful offensive model the same week it announced a defensive fund, calling it “the imbalance that we’re talking about, in a single news cycle.” He argued that every public benchmark the labs compete on is an offensive one, and that if defense is now the priority, it should be measurable.

The $1 Billion Subsidy and the Asymmetry Problem

OpenAI committed $1 billion in subsidized access, training, and technical support through its Daybreak for Frontline Defenders initiative, targeting water systems, electric grid operators, small banks, and state and local governments. The initiative pairs model access with a pilot through the Multi-State Information Sharing and Analysis Center to help a group of essential-services providers validate findings and coordinate remediation.

The criticism from defenders focuses on capacity. Mike Bell, founder and CEO of Suzu Labs, noted that a small water utility or county health department receiving Daybreak access often lacks staff who can interpret what the model surfaces, prioritize what matters, or execute a fix without breaking something else. Shipping model credits alone, he argued, gives someone a tool they lack the staff to use; pairing those credits with a forward-deployed engineer would produce real change.

The comparison table below places the competing approaches side by side, with each figure drawn from a named source.

Approach Reported result Source
Standalone SAST rule engines Roughly 35.7% precision rate GitHub / TechTimes
LLM reasoning combined with SAST Up to 91% false-positive reduction GitHub / TechTimes
GPT-5.6-Cyber (OpenAI benchmark) 95% of advanced cybersecurity requests, vs 1.5% for general GPT-5.6 Sol SiliconANGLE
Astra on ExploitBench 100% vs 78.5% for GPT-5.6 Sol The Hacker News
Disclosed vulnerabilities (Action1, 2025 vs 2024) +92% overall, +128% for RCE flaws BleepingComputer

The benchmark numbers carry a caveat worth stating plainly: the 95 percent and 100 percent figures are vendor-reported results on vendor-selected benchmarks. They have not been independently reproduced. The trend, that specialized cyber models outperform general-purpose models by a wide margin, is credible. The exact percentages are OpenAI’s own measurement.

An Audit Checklist for AI-Assisted Defense

Teams using these tools should treat their output as suggestions that require verification, not as unquestionable answers. Each item below corresponds to a specific failure mode observed in 2026.

  • Gate every AI finding behind a human. Cloudflare’s service does not let a rule or patch take effect without explicit approval. Replicate that: no auto-remediation without a named approver.
  • Validate before you rate. Cloudflare validates each finding before assigning a risk rating, then raises the rating only on production evidence like traffic or active probing. Do not rank raw scanner output.
  • Watch for self-rewriting artifacts. PROMPTFLUX rewrites itself hourly. Detection must key on behavior and intent, not hashes or static signatures.
  • Assume false positives are the default. A 35.7 percent precision baseline means two of every three flags may be noise. Budget analyst time accordingly and track your own precision rate.
  • Do not depend on NVD alone. With roughly 30,000 pre-March 2026 CVEs reclassified as “Not Scheduled,” correlate across vendor advisories, CISA’s KEV catalog, and your own asset inventory.
  • Measure defense, not just offense. If you run models, log the share of effort spent on defensive versus offensive tasks, and hold the number up against the lab’s own claims.
  • Pair credits with people. Model access without staff to interpret output is a gift that goes unused. The $1 billion subsidy only lands if the recipients can act on what the model surfaces.

The defining fact of 2026 is that discovery and exploitation now move at the same speed, because they are powered by the same models. Organizations that stay ahead will treat every AI finding as a hypothesis to verify, every patch as a decision a human owns, and every signature as a control that a self-rewriting adversary has already learned to outrun.

More in-depth coverage from this blog on closely related topics:

Sources and References

Sources cited while researching and writing this article:

Rafael

Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...