AI Patch Generation Fails Security Test
TL;DR. New research shows AI models like ChatGPT and Claude frequently fail to fix cybersecurity bugs, often introducing fresh code vulnerabilities. - 1Password researchers tested OpenAI's ChatGPT 5.5 and Anthropic's Claude Opus 4.8 on six complex CVEs. - The AI models achieved a success rate below 47%, often only addressing subsets of vulnerabilities. - Researchers noted AI-generated patches frequently introduced new bugs or altered application behavior. - Other industry reports corroborate findings on AI code generation's low security pass rates.
- AI-generated code patches from models like ChatGPT and Claude often fail to fix vulnerabilities completely.
- More than half of AI-generated patches were found to introduce new security flaws or modify application behavior.
- Research by 1Password tested high-impact CVEs, finding a success rate of under 47% for full remediation.
- Similar reports from Veracode confirm that AI models struggle with security in code generation.
Sources
- More than half of AI-generated patches are broken — cyberscoop.com