
Furthermore, once one of these banned instructions make it into the context window, the whole session is poisoned, and the model will often refuse to continue without manual intervention. The researchers ran many tests to find short strings that would trigger model safety guardrails reliably, but it’s worth noting that the identified strings were different between the tested models: Claude Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek V4 Pro, and Kimi K2.6.
During baseline tests the AI agents managed on average to obtain full account admin in 54% of the 154 attack runs and full compromise (admin + persistence) in 36% of tests. With the context bombs in place, their success rate dropped to 5% for admin access and 1% for full compromise. Also, in 91% of baseline attack runs, the agents managed to complete at least one of ten possible attacks paths, but their average success rate dropped to 15% with the context bombs.
The models from Western AI labs — Opus and Gemini — proved the most capable at reaching full admin access, with 93% and 70% success rates, but were also the most impacted by the context bombs with both their success rates dropping to 0%. This shows that the content safety guardrails are much stronger in these models compared to the Chinese ones that were tested.
