CISOOnline

AI can find zero-days but still can’t reliably write secure code

Such harnesses can retrieve relevant files and architectural documentation, provide threat model information, list approved coding patterns, run compilers and tests, invoke static and dynamic security testing tools, and enforce approval gates that prevent AI agents from continuing until identified failures are addressed.

For example, OpenAI partnered with Trail of Bits to use its models to find vulnerabilities in open-source projects critical to internet infrastructure and help develop patches for them. For the project, dubbed Patch the Planet, researchers built specialized workflows and harnesses, and, as of Aug. 11, the initiative listed 1,250 reported issues across 49 codebases, 271 authored fixes, and 146 patches accepted upstream.

Xint has observed a similar effect in internal testing, according to Kwak. The company evaluated frontier models against 208,000 lines of code containing 17 injected vulnerabilities. Bare prompted loops found between zero and one of the flaws, whereas models operating inside Xint’s specialized harness found between 11 and 14.



Source link