
“The attack operates under a chosen plaintext threat model, which is the most common assumption used for studying ciphers like AES,” it explained in the blog post. This assumes that an attacker is able to request that the defender encrypt arbitrary inputs with a fixed, unknown key, and then gets to see the corresponding output. In this case, it assumes the attacker can request the encryption of 2^105 (about 4 billion billion billion) chosen plaintexts. “This attack is therefore completely impractical but quantifies the attack cost against AES under these assumptions,” it said.
Claude Mythos Preview discovered the attack almost entirely autonomously, Anthropic said, adding, “A researcher at Anthropic built a scaffold that enabled Claude to pose hypotheses, run experiments to experimentally validate or refute these hypotheses, and then asked Claude to design an attack that improves on the best cryptanalysis of AES.”
Even with AI accelerating some parts of the research, more time is spent verifying the correctness of results. The improved attack on AES was discovered in about a week, but took two researchers nearly a month to validate, Anthropic said.
