CyberDefenseMagazine

Frontier AI and the Vulnerability Gap in OT


The landscape of cyber defense and offense is shifting beneath our feet. Over the past few weeks, the release of “frontier” AI models has accelerated the arms race, forcing a necessary re-evaluation of our current security strategies. While these models promise to revolutionize how we find and fix flaws, they also introduce a significant “velocity problem” that is particularly acute for Operational Technology (OT) and Internet of Things (IoT) environments.

Two major announcements have recently set the pace for this transition; Mythos and Daybreak. Anthropic’s Claude Mythos Preview was launched on April 7th. This model is paired with “Project Glasswing,” an initiative aimed at giving defenders early access to stay ahead of the “avalanche” of vulnerabilities Mythos is capable of uncovering. OpenAI’s Daybreak was announced on May 11th. Daybreak ups the ante by not only discovering vulnerabilities but also generating patches for them using the OpenAI Codex engine.

While there is some debate as to whether these frontier models can find vulnerabilities that human researchers cannot, it is clear that the volume and velocity of finding vulnerabilities, even complex ones, is orders of magnitude larger than before these frontier models. For operators of OT systems, the most devastating aspect is that AI can be both operating system (OS) and application agnostic. With IT systems there are a handful of OS’s used; for OT and IoT there are over 150,000 of them.

While OpenAI emphasizes a “human-in-the-loop” approach, this may be a flawed safety net. The speed at which these models operate can overwhelm humans (remember there is a serious shortage of cybersecurity talent). In addition, Daybreak is designed to handle tasks beyond human manual capability, raising fundamental questions about what meaningful human oversight actually looks like. Furthermore, Daybreak’s relatively open access contrasts with Mythos’s more controlled, “closed” distribution—a distinction that raises valid safety concerns in these early stages of AI-driven security.

It’s the “upping the ante” by OpenAI so quickly (a few weeks) after Mythos came out that should also worry security teams. Did OpenAI make a bigger, bolder, more extensive release that many can access because it was safe, or was it to match Anthropic from a marketing perspective?

The Speed Trap and the OT Reality Check

It is important to recognize that these models don’t fundamentally change the nature of cybersecurity; they change its velocity. We have long lived in a world where vulnerabilities are discovered faster than they can be patched. AI simply widens that gap.

This “churn” is magnified in the worlds of OT and IoT. While AI-generated patches might work well for generic, open-source software, the custom and unique nature of OT systems presents a massive hurdle. Consider the scale: while IT relies on a handful of standard operating systems, the OT and IoT ecosystem is fragmented into over 150,000 different operating systems. The other characteristic of OT and IoT systems is the tight coupling of devices and the applications that manage them. For an approach like Daybreak to work requires deep and context-sensitive understanding of that device-to-application integration, otherwise the patches generated on the fly could render a mission critical system inoperative.

The Missing Piece: Emulation and Testing

The concept of automated patch generation “on the fly” isn’t entirely new. Programs like ARPA-H’s UPGRADE are pursuing similar goals for medical systems. However, there is a critical difference: UPGRADE focuses on creating detailed emulators to test patches before they are deployed.

OpenAI’s Daybreak currently lacks this critical emulator step. In an OT environment—where a bad patch can lead to physical downtime or safety hazards—deploying AI-generated code without rigorous, environment-specific testing is a non-starter. Compounding this issue is the fact that neither Mythos nor Daybreak currently have direct OT partners to bridge this domain-knowledge gap.

Confidence vs. Competence

Early experiences with AI code generation I’ve personally used such as using agents like Manus and Claude, reveal a recurring pattern: the AI is often “brimming with confidence” but lacks the nuanced understanding to execute complex tasks flawlessly.

Here is my experience using AI for managing website pages. When faced with a roadblock, AI agents tend to “hallucinate” solutions—creating a chain of new code (like endless WordPress plugins) to fix issues created by previous iterations. In these early days, relying on AI for automated patch code generation risks creates a “steaming mass of technical debt and complexity” rather than a secure system. This is a serious danger of automated code generation within the echo chamber of a frontier AI model; it can create complexity that even the device manufacturer will have trouble sorting out.

Moving Toward “Targeted Autonomous Remediation”

We must adapt our objectives to match the current maturity of these tools. The industry is on a path that starts with automated Vulnerability Assessment and Risk-Based Prioritization, moving toward Targeted Autonomous Exposure Remediation. AI is also following that path but it does not mean it is here today, safe, or the best solution. Organizations should be getting their hands dirty with all of these (especially autonomous remediation) in order to properly assess AI for this, but it should be with significant human oversight and in lab (not production) environments.

Until then, we must remember that Defense in Depth remains more important than ever. Patching is just one piece of the puzzle. Foundational controls—including managing credentials, updating vulnerable software, and non-human identity management—must work in tandem to provide compensatory protection when patching gaps exist.

Ultimately, exposure management is a balancing act between operational needs, opportunities, and cybersecurity risks. AI is a powerful tool to increase our velocity, but in the unique world of OT, speed without precision is just another vulnerability.

About the Author

John Gallagher is the VP of Viakoo Labs at Viakoo Inc. He is head of Viakoo Labs, a research unit focused on OT (operational technology) vulnerability remediation and best practices. John is a member of SIA’s Cybersecurity Advisory Board and has presented at industry events including ISC West, ISC East, GSX, HIMSS, and is a frequent blogger and webinar host. John has BS in Engineering and MBA from Cornell University and has lived and worked in Europe and Asia and speaks (basic) Japanese.

John can be reached online at [email protected] or on LinkedIn and at our company website https://www.viakoo.com/



Source link