Tldrsec

[tl;dr sec] #349 – Vulns & Exploits in the AI Era, Package Manager Sandboxing, Testing AI Sandboxes


I hope you’ve been doing well!

One of everything

Look at this stuff, isn’t it neat? Wouldn’t you think my collection’s complete?

Wouldn’t you think I’m a hacker who has… everything?

It do be like that right now Happy hunting my friend!

The task ends. Your AI agent’s access shouldn’t linger.

Opal Zero evaluates access requests against policy and context, automates routine approvals, and escalates sensitive requests to humans. Permissions are scoped, time-bound, and enforced through the MCP gateway you already run.

Developed with design partners Databricks, Faire, Elastic, and Superhuman.

Nice, love the least privilege focus and automatically removing permissions after a task is completed  

AppSec

How XBOW Finds IDORs with High Accuracy and Low Noise
XBOW’s Alvaro Muñoz writes about why DAST tools miss IDORs, which are design flaws rather than implementation bugs. You can’t grep for them, and a scanner that sees a valid session, a well-formed request, and a 200 response marks it a pass, because it has no idea which data belongs to which user. Their first build had the same problem, with one agent hunting for IDORs and a second arguing both sides of each finding before ruling on it, neither of them knowing how the application was meant to behave. So they had agents log in as each role and map what it can and cannot do before any testing starts. That map is what a finding gets checked against, so the judgment rests on what the role was observed doing rather than what the model assumes should be allowed.

HTTP/3 in Burp Suite – it’s time to find a bigger wordlist
PortSwigger’s Tom Stacey added HTTP/3 support to Turbo Intruder, and it can now send over 100,000 requests per second from a laptop on Wi-Fi, more than three times what it managed over HTTP/1.1. Running it from a cloud server in the same region as the target pushed his best result to about 180,000. Burp Suite Professional users also get a new AUTO engine that picks the newest HTTP version the target supports and adjusts its own settings during the attack.

On the testing side, the HTTP/3 engine adds two new race condition techniques from recent research, the Single Datagram Attack and a method built on QPACK blocked streams, which line requests up more tightly than the Single-Packet attack and can catch smaller race windows. Turbo Intruder can also test HTTP/3 downgrade attacks, where a line break hidden inside a header value splits into a new header, like Transfer-Encoding, once the server converts the request to HTTP/1. A separate HTTP/3 Adapter extension brings HTTP/3 to the rest of Burp, so testers can reach sites that only speak HTTP/3.

Most AI security advice stops at prompt injection. This guide breaks the problem into the three layers that actually get attacked: the infrastructure that hosts and runs your models, the software and data the application depends on, and the entry points and business logic a user can reach.

Each layer gets concrete steps, the failure modes that show up once an application is live, and what to monitor so you catch problems in production rather than in a postmortem.

That layer breakdown makes sense to me, those are all indeed important. Curious about the tips for each attack surface  

Supply Chain

Graphalgo campaign spreads to Terraform providers and Go Modules
Aikido’s Oliver Smith found malware in two Terraform providers (gocommunity-io/dockerd and kreuzwenker/docker) and two Go modules, the first time Aikido has seen malware distributed through Terraform providers. The code is a Go port of the Graphalgo NPM malware and shares that campaign’s Slack and blockchain infrastructure. To make the packages look legitimate, the threat actor typosquatted the 56-million-download kreuzwerker/docker provider, set up fake Go package ecosystems at gocommunity.io and gogets.dev, and backdated commits on one module so pkg.go.dev showed a November 2025 release date.

Once installed, the providers stay dormant unless the SHA256 hash of two specific Terraform variables matches a hardcoded value, so they look clean everywhere except the environment the attacker is after. When the hash matches, they decrypt and run a Go RAT that takes encrypted commands over both Slack and an Ethereum smart contract on Arbitrum Sepolia, and each infected host uses its own ephemeral key pair so no victim can read another’s traffic. The attacker’s Slack channel shows only 18 infected hosts since July 2026, a small number that fits a targeted operation. Oliver adds that Terraform users make especially valuable targets, since they deploy infrastructure and often have a direct path to production credentials.

Honestly, pretty thoughtful targeting and OPSEC  

Package Manager Sandboxing
Andrew Nesbitt looks at how package managers handle install-time code, which comes down to three options: gate it behind an allowlist, confine it in a sandbox, or replace it with data the manager reads instead of runs. Most package managers have proposed sandboxing, but few have shipped it. Homebrew 7.0.0 splits installation into a networked fetch phase and an offline install phase, blocks builds from reading your home directory, and swaps arbitrary Ruby in post_install for declarative signed steps. opam has sandboxed since 2018, Swift Package Manager’s confinement silently does nothing on Linux, and Nix and Bazel isolate builds for reproducibility rather than security.

Everywhere else it’s open issues. pnpm, Deno, RubyGems, and Spack all propose Landlock on Linux and Seatbelt on macOS, the same stack Claude Code and Codex adopted. Cargo’s sandbox issue has been open since 2018 and Poetry closed the request as out of scope. Removal has worked better than confinement, with Go and Elm never shipping hooks at all and NuGet deleting install.ps1 back in 2017. All three options still leave the same gap, since something has to validate what the sandbox hands back before the package manager acts on it with full privileges.

Fantastic survey of how various package managers handle sandboxing. Very dense and to the point, tons of supporting links, love it!

Blue Team

Where do detection ideas come from
Gary Katz and Jason Deyalsingh write about where detection engineering ideas should come from, now that AI has made writing the rules themselves the easy part. They describe four information sources that can be a starting point or a filter mechanism for prioritization: visibility (what logs you actually have), existing detections (open-source repos like SIGMA that AI can triage for your team), environment (what you’re protecting and where your EDR has blind spots), and threat intelligence.

On threat intelligence, they split it into strategic (where to focus), operational (which adversaries and TTPs to prioritize), and tactical (how attacks are executed), but note that most tactical intel focuses on atomic indicators like IPs, domains, and hashes rather than the procedural depth you need for a solid detection, with red-team articles and resources like Tired-Labs filling that gap. They say chasing every piece of tactical intel without filtering through the other three sources leads to burnout, and teams get more out of using strategic and operational intel to decide what’s worth building.

Vulnerability Discovery and Exploitation Trends in the AI Era
Google Threat Intelligence Group’s Robin Grunewald, Supriya Mazumdar, and Kelli Vanderlee analyzed vulnerability disclosure and exploitation trends from January 2025 to August 2026, finding that monthly CVE disclosures doubled from 5,045 to 10,740, while in-the-wild exploitation rose from 10.5 to 18 per month and zero-day exploitation grew more modestly from 8 to 11 per month. Exploitation growth (+127%) closely mirrored disclosure growth (+128%) from May 2026 onward, suggesting attackers are weaponizing n-days rather than discovering new zero-days, though only 0.23% of disclosed CVEs are ever exploited.

GTIG found that AI-discovered vulnerabilities tend to be higher impact (58% Medium Risk vs 28% for non AI CVEs), and 50% result in RCE, which GTIG attributes to researchers pointing agents at critical infrastructure and privilege boundaries rather than running broad scans. GTIG also counted 2,076 AI-related CVEs, half of them in agent orchestration frameworks like Flowise and Langflow, and several are already exploited in the wild, including command injection in LiteLLM and unauthenticated RCE in Langflow.

It’s interesting how few CVEs are ever exploited, and how it seems like attackers are weaponizing n-days over zero-days. I will say models are very great at creating exploits from patches or even just general hints about a vulnerability. Regarding AI-discovered vulns being higher impact, I think that’s just because people are sitting on thousands of vulns and why bother reporting low sev ones when you have more RCEs than you have time to look through. Errr speaking for a friend  

AI + Security

$1 Million Sandbox Challenge Uncovers Linux Kernel Flaws
Vercel ran a two-week, $1M bug bounty program challenging researchers to escape their Firecracker-based microVM sandbox, receiving 1,285 reports with ~$325k in payouts across 1 Critical, 7 High, 15 Medium, and 49 Low severity findings. To handle the massive report volume, Vercel built an agentic triage system using their Eve framework running on Kimi K3 that reads reports, checks policy compliance, hunts duplicates, pulls source code, and executes PoCs in real sandboxes, which they plan to open source. Trail of Bits white-box testing yielded 20 findings and key architectural recommendations, including “stop trusting the guest” by deriving all values crossing the VM boundary server-side or signing them with unreachable keys.

Paulos Yibelo claimed a full VM escape zero-day, which was then confirmed by Vercel CEO Guillermo Rauch as a KVM zero-day.

Investing in the common tech that many rely on, love to see it  

Escaping SPACE: Part I
Perplexity tested whether AI agents with root access inside Firecracker microVMs could escape to the host or bypass network restrictions, evaluating nine models across 216 runs with varying source code access and network permissions. No model achieved VM-to-host escape in any of the 108 sandbox escape attempts, but once the policy allowed package repositories through, four models (Claude Opus 5.0, GPT-5.6 Cyber, GPT-5.6 Sol, and Kimi K3) bypassed network confinement in 11 of 54 partial-network runs, exploiting the same gap between domain-based policy and IP-based enforcement two different ways.

The models bypassed network restrictions through DNS spoofing (forging responses to map allowed domains like pypi.org to blocked IPs) and IP-sharing attacks (exploiting CDN infrastructure where services like Fastly Fiddle or Taboola shared IPs with allowed domains). After implementing source-address validation via nftables and TLS-terminating proxies that verify SNI and HTTP Host headers, none of the models succeeded. They then ran the network policy tests against ten third-party sandbox platforms and found the same bypass on eight, including E2B, Vercel Sandbox, Modal Sandboxes, and Deno Sandbox.

Great description of the experimental design and motivation (and diagrams), and nice of them to test other sandbox platforms and report their findings. Also, I want to point out that it’s smart to give the agent the source code to the system containing it- similar to how you’re going to get better results from a white box pen test, you’re better testing your sandboxing if the agents understand how it’s implemented.

Misc

Wrapping Up

Have questions, comments, or feedback? Just reply directly, I’d love to hear from you.

If you find this newsletter useful and know other people who would too, I’d really appreciate if you’d forward it to them

P.S. Feel free to connect with me on LinkedIn  



Source link