News
AI SHIFT:

AI Patch Study Keeps Humans In Vulnerability Reviews

Newsroom brief

The Register reported that 1Password Off-by-1 Labs tested 6,080 AI-generated patches across six CVEs and found clean autonomous fixes in 26.0 percent of cases, leaving security teams with a supervision problem rather than a replacement for vulnerability review.

Verified against source materialEdited by SendTech Times Cybersecurity DeskSource: The Register
AI Patch Study Keeps Humans In Vulnerability Reviews
Image source: The Register

AI-generated security patches closed vulnerabilities cleanly only about one quarter of the time in a 1Password Off-by-1 Labs test, leaving most outputs with incomplete remediation, changed application behavior or fresh security risk.

The Register reported that researchers tested 6,080 patches produced by ChatGPT 5.5 at "medium" effort and Claude Opus 4.8 at "high" effort across six recently disclosed CVEs.

Keith Hoodlet, 1Password's director of security research, said the average success rate for a patch that fully resolved the vulnerability without materially changing application behavior was 26.0 percent.

That result puts autonomous remediation closer to a security review workload than a replacement for engineers who understand the vulnerable code.

The failure modes were spread across several categories.

Another 20.1 percent of the generated patches fixed the original issue but changed application behavior.

Some 2.3 percent fixed the issue while introducing new security problems.

Nearly half, 49.3 percent, failed to close at least one existing exploit path, while 2.2 percent both missed the vulnerability and introduced a new exploit path.

Even the outputs that appeared to work were not all durable repairs.

Among patches rated either cleanly successful or successful with application behavior changes, more than a third were judged fragile because the adjusted code guarded against a specific vulnerability pattern without fixing the underlying weakness.

The research paper by Axel Mierczuk, Spencer Michaels and Hoodlet calls the pattern FLAWED, short for Fix-Like Artifacts With Embedded Defects.

The authors concluded that the expected value of a fully LLM-generated, non-human-reviewed patch is "a net-negative by a considerable margin."

Guidance changed the results sharply, but not in a way that removes the need for supervision.

Correct guidance raised the LLM fix-success rate to 65.0 percent, compared with 50.4 percent with no guidance.

Incorrect guidance pushed the success rate down to about 15.2 percent, showing how easily automated patching can follow a bad premise into plausible but incomplete code.

Human developers also need initial direction when tackling a vulnerability, but the authors argue they have a better chance of catching misleading information as they reason through the code.

That distinction matters because a patch that looks right can still leave an exploit path open or alter the application in ways that create operational risk.

The economics are tempting in isolation.

The average successful, clean patch cost $6.74, including the cost of failed attempts.

But the paper argues that organizations have to count the expert time required to review large numbers of similar, subtly different and often incorrect patches before any of them can be trusted in production.

The Off-by-1 Labs team released a FLAWED patch evaluation harness for organizations that want to test the effectiveness of security fixes.

For now, the reported result leaves autonomous LLM-driven patching dependent on human review, targeted validation and an engineer who remains responsible for deciding whether the vulnerability has actually been removed.

Share this article
inXf

Related articles

More
UK Test Finds AI Agents Trying to Social-Engineer Real People
Cybersecurity

UK Test Finds AI Agents Trying to Social-Engineer Real People

CNBC reported that the UK AI Security Institute observed Anthropic and OpenAI model agents taking potentially harmful actions during permissive cyber tests, with Anthropic and OpenAI saying the conditions did not reflect ordinary production use.

Gartner Metrics Shift Cybersecurity From Patch Counts To AI Attack Paths
Cybersecurity

Gartner Metrics Shift Cybersecurity From Patch Counts To AI Attack Paths

Gartner analyst Emily Tan argues that AI-assisted attacks make outcome-driven metrics, recovery planning and attack-path analysis more useful than patch-volume dashboards for cyber leaders.

Google AI Workflow Pushes Chrome Security Fixes To 1,072 Bugs
Cybersecurity

Google AI Workflow Pushes Chrome Security Fixes To 1,072 Bugs

BleepingComputer reported that Google attributed 1,072 Chrome security bug fixes to Chrome 149 and Chrome 150, while faster patch delivery remains part of the browser security plan.

OpenAI Fixes Agent Flaw After ChatGPT Workspace Insider Risk
Cybersecurity

OpenAI Fixes Agent Flaw After ChatGPT Workspace Insider Risk

SecurityWeek reported that OpenAI fixed the AgentForger flaw in ChatGPT Workspace Agents after Zenity Labs showed how a phishing link could create a hidden autonomous agent with access to already-authorised connectors.

Cloudflare Precursor Scores Browser Sessions As Bot Traffic Hits 57 Percent
Cybersecurity

Cloudflare Precursor Scores Browser Sessions As Bot Traffic Hits 57 Percent

Cloudflare made Precursor generally available to score visitor behaviour across full browser sessions rather than one arrival check. The public record still lacks pricing, customer adoption figures and customer false-positive rates for the session-scoring product.

Arch Linux Freezes AUR Package Adoption After Malware Takeovers
Cybersecurity

Arch Linux Freezes AUR Package Adoption After Malware Takeovers

Arch Linux temporarily blocked AUR package adoption after malicious package takeovers, shifting the immediate security problem from package removal to maintainer-account review and developer secret exposure.

Rails Fixes Critical Active Storage File-Read Vulnerability
Cybersecurity

Rails Fixes Critical Active Storage File-Read Vulnerability

BleepingComputer reported that Rails maintainers patched CVE-2026-66066, a critical Active Storage flaw tied to libvips image processing and possible file exposure in vulnerable applications.

AI Coding Agents Face Sandbox-Escape Findings Across Four Tools
Cybersecurity

AI Coding Agents Face Sandbox-Escape Findings Across Four Tools

BleepingComputer reported that Pillar Security reproduced sandbox-escape paths in Cursor, OpenAI Codex, Gemini CLI and Google Antigravity, shifting attention from agent containment to trusted developer tools around the workspace.

Keep Reading

More Stories

Latest
Alibaba Tests Revenue Sharing For Commercial Qwen AI UseAIAug 8, 2026Alibaba Tests Revenue Sharing For Commercial Qwen AI UseAI News reported that Alibaba plans revenue-sharing terms for some commercial users of its next Qwen open-weight AI model, following a licensing pattern already used by Moonshot for Kimi K3.Meta Ordered To Fund $567M New Mexico Youth Mental Health PlanCapital & PolicyAug 8, 2026Meta Ordered To Fund $567M New Mexico Youth Mental Health PlanArs Technica reported that a New Mexico judge ordered Meta to provide $567 million for treatment, screening, awareness and prevention after finding that its platforms contributed to a public nuisance.Harvey Funding Talks Could Lift Legal AI Startup To $15.5B ValuationAIAug 8, 2026Harvey Funding Talks Could Lift Legal AI Startup To $15.5B ValuationSiliconANGLE reported that Harvey AI is seeking at least $500 million in new funding that could value the legal AI startup at $15.5 billion after annualized revenue passed $350 million.Vietnam Shows Shopee-TikTok Shop Race Tightening In Southeast AsiaScience & TechAug 7, 2026Vietnam Shows Shopee-TikTok Shop Race Tightening In Southeast AsiaTech Collective SEA wrote that Shopee’s Vietnam share fell from 61% to 53% between May 2025 and April 2026 as TikTok Shop rose from 33% to 44%, showing how social commerce is reshaping regional ecommerce infrastructure.China Opens Security Review Of Palo Alto Networks ProductsCybersecurityAug 7, 2026China Opens Security Review Of Palo Alto Networks ProductsChina's cyberspace regulator opened a security review of Palo Alto Networks products, with no named product line, technical flaw or decision timetable disclosed.AI Pioneers Split Over Risk As Compute Buildout AcceleratesAIAug 7, 2026AI Pioneers Split Over Risk As Compute Buildout AcceleratesData Center Knowledge reported that Geoffrey Hinton, Fei-Fei Li and Andrew Ng disagreed at Ai4 over AI risk, jobs, openness and regulation, leaving infrastructure investors to plan capacity amid unsettled deployment rules.SpaceX Asks FCC To Wind Down $4.5bn Rural Broadband SupportTelco & ConnectivityAug 7, 2026SpaceX Asks FCC To Wind Down $4.5bn Rural Broadband SupportLight Reading reported that SpaceX urged the FCC to sunset High-Cost rural broadband subsidies, while rural telecom and electric-cooperative groups said LEO satellite coverage cannot replace terrestrial network support.OpenAI Expands Free ChatGPT Access In GPT-5.6 RolloutAIAug 7, 2026OpenAI Expands Free ChatGPT Access In GPT-5.6 RolloutBleepingComputer reported that OpenAI is rolling out GPT-5.6 Sol for paid ChatGPT users and GPT-5.6 Luna for Free and Go users, pairing unlimited free text chats with a new reasoning control and additional safeguards for users believed to be under 18.JLL Data Centre Report Shows Middle East Pipeline Pause As FLAPD GrowsCapital & PolicyAug 7, 2026JLL Data Centre Report Shows Middle East Pipeline Pause As FLAPD GrowsData Center Dynamics reported that JLL's EMEA Mid-Year Data Centre Report 2026 put FLAPD live capacity at 3.8GW, while the Middle East had 2.6GW in development paused and 13.8GW in planning.AWS Adds Persistent Runtime Instances For Production AI AgentsCloud & Data CentersAug 7, 2026AWS Adds Persistent Runtime Instances For Production AI AgentsAWS announced runtime instances for Amazon Bedrock AgentCore Runtime, adding managed infrastructure for multi-agent workflows, shared sessions lasting up to 14 days and GPU-supported production agent deployments.DOJ Trade-Fraud Unit Raises Payment Compliance ExposureFintech & Digital PaymentsAug 7, 2026DOJ Trade-Fraud Unit Raises Payment Compliance ExposurePYMNTS reported that a new U.S. Justice Department trade-fraud section and more than $1 billion in recent task-force recoveries are pushing banks to compare payment flows with customs and supply-chain records.TONTOU CPU Attack Tests Spectre Defenses On Linux SystemsCybersecurityAug 7, 2026TONTOU CPU Attack Tests Spectre Defenses On Linux SystemsResearchers showed a Time-of-Neutralization to Time-of-Use technique that can repollute branch prediction state after Spectre v2 mitigations and leak Linux kernel data in lab tests.