An AI model found mathematical flaws that two years of expert cryptographic review missed. It did not break your encryption, and the headlines saying otherwise are wrong. The real story is stranger, and the most useful part of it has almost nothing to do with cryptography.
Before it succeeded, the model told its own researcher the problem was unsolvable. It was articulate about why. It was also completely wrong, and the fix was not a better model—it was a better scaffold and a researcher who refused to accept the first answer.
This is the most widely shared AI story of the week, and it is being flattened in transit. The version circulating on social media is that AI has started breaking encryption. That is false. The version in the serious trade press is careful and accurate but stops at the caveats.
Underneath both is a set of details that I think matter more than the cryptography — because they describe, in the hardest technical domain in computer science, exactly the working pattern I keep arguing for in ordinary businesses.
What Anthropic Actually Announced
Anthropic published research reporting that Claude Mythos Preview — an unreleased model above the Opus tier — discovered improved attacks against two cryptographic algorithms.
Finding One: HAWK
HAWK is a post-quantum digital signature scheme, a third-round candidate in NIST's competition to select cryptography that will survive quantum computers. It had been through roughly two years of expert human review.
Claude identified a previously undetected symmetry in the mathematical lattice HAWK's security depends on, and used it to build a faster key-recovery attack. The effect is severe in relative terms: the scheme's effective key strength is cut in half. For the smallest HAWK-256 configuration, the estimated cost of recovering a key falls from 2 to the power of 64 operations to 2 to the power of 38.
In practice, HAWK's key sizes would need to double to restore the intended security level — which, as CyberScoop reports, erases much of what made it an attractive candidate in the first place. Anthropic is explicit that the finding is specific to HAWK and does not extend to other NIST post-quantum candidates or to lattice-based cryptography generally.
Finding Two: The Möbius Bridge
The second result targets seven-round AES-128. Full AES-128 uses ten rounds; researchers routinely study reduced-round variants to measure how much security margin a cipher has in reserve.
Building on established meet-in-the-middle methods, Claude produced a fingerprinting technique Anthropic named the Möbius Bridge. It removed a guessing step that had required checking 256 values, making the best known theoretical attack on seven-round AES 200 to 800 times faster depending on how you measure.
The practicality caveat is enormous and Anthropic states it plainly. The attack operates under a chosen-plaintext model requiring roughly 2 to the power of 105 chosen plaintexts — on the order of 400 octillion messages. Anthropic's own description is that the attack is completely impractical and serves to quantify attack cost rather than to threaten anything.
What This Does Not Mean
Because this is the part that gets lost, let me be unambiguous:
- No production system is affected. Not one.
- Full ten-round AES-128 — the cipher protecting essentially all data in transit — is untouched.
- HAWK is not deployed anywhere. It is a candidate under evaluation, which is precisely the stage at which you want flaws found.
- Nothing here suggests encryption is broken, weakening generally, or about to be.
Anthropic followed responsible disclosure throughout: notifying HAWK's designers in June, coordinating public release through a NIST mailing list, and briefing government and industry partners before publishing.
An industry view worth quoting sits in CyberScoop's coverage. Ellen Boehm of Keyfactor argued the research demonstrates the NIST evaluation process working as intended — a candidate scheme was stress-tested and found wanting before deployment rather than after. That is the system succeeding, not failing.
A flaw found in a candidate algorithm during evaluation is a good outcome. The alarming version of this story is the one where the same capability is pointed at something already deployed, and nobody publishes.
The Detail Everyone Skipped: It Refused
Here is where the story stops being about cryptography.
Anthropic published the actual prompts its researchers used, typos included. Reading them is more instructive than the results.
On the AES problem, Claude initially would not engage. It assessed the task as impossible and said so — that AES-128 at five, six and seven rounds is genuinely hard, that it had found nothing because there is nothing easy to find, that this is the most-studied block cipher in existence. It suggested that if the researcher wanted a different outcome, the target would have to change.
Every one of those statements is defensible. AES really is the most-studied block cipher in existence. A confident, well-reasoned, articulate assessment — and wrong.
The researcher's fix, written in a note Anthropic published verbatim, was that the models tend to think it is impossible to solve so they do not try, and they need a good amount of prompting.
That is the whole lesson, and it generalises far outside cryptography.
A model's assessment that something cannot be done is not evidence that it cannot be done. It is a prediction generated from the same training distribution that produced its other outputs, and it carries the same error rate. Yet in practice most people treat a refusal or an it-is-not-possible as authoritative in a way they would never treat a confident factual claim.
I would guess a meaningful share of the automation projects abandoned in the last two years died at exactly this step. Somebody asked, the model said no, and nobody pushed. The task was feasible. The framing was not.
The HAWK Version of the Same Thing
The pattern repeated on the other problem. According to The Decoder's account, the HAWK work ran as a multi-agent system in which one agent initially dismissed the promising idea as infeasible — and a second agent found the way to exploit it fully.
One instance said no. Another instance, given the same idea, said yes and was right.
If that does not make you re-examine how much weight you place on a single model response in your own workflows, I am not sure what would.
The Human Was Not a Cryptographer
The next detail is the one I expect to be quoting for the rest of the year.
The Anthropic researcher who worked with Claude on the HAWK attack had a background in theoretical computer science but was not an expert in lattice-based cryptography. The Decoder's summary of his contribution: his role was mostly limited to project management.
Sit with that. A finding that two years of specialist review missed was produced by a capable generalist directing a model, in a domain the human did not specialise in.
This is the Bending Spoons pattern arriving in the most technically demanding field it could possibly arrive in. Humans doing specification, framing and judgment. Machines doing production. The human contribution was knowing which questions were worth asking and refusing to accept the first answer — not out-computing a lattice.
It is also the clearest counterexample I have seen to the idea that AI leverage requires deep domain expertise to unlock. What it required here was research taste, a well-built harness and persistence.
The scarce input was not cryptographic expertise. It was someone who knew how to structure a problem, build the loop, and keep going after the model said it was impossible.
The Scaffold Was the Breakthrough
On the AES result, Anthropic is specific about the mechanism: a researcher built a scaffold that let Claude pose hypotheses, run experiments to validate or refute them, and then design an attack improving on the best known cryptanalysis. The model then discovered the attack almost entirely autonomously.
Note the order of operations. The model alone, asked directly, refused. The model inside a hypothesis-and-experiment loop, with a human insisting the problem was worth attacking, produced novel mathematics.
Same model. Same underlying capability. Entirely different outcome, determined by the structure wrapped around it.
This is the single most important thing in the story for anyone running a business, and it is the argument I have been making since I started writing here. Most organisations do not have an AI problem, they have a systems problem. They have access to the same models as everyone else and get worse results, because the model is not where the leverage lives.
Anthropic's researchers did not have a better Claude than the one that refused. They had a better loop.
Hundreds of hours of human validation followed the discovery, which is worth noting for balance — this was not a fire-and-forget result. But the discovery itself came from the harness.
What $100,000 Buys Now
Each of the two results cost roughly $100,000 in API usage. The HAWK work ran about sixty hours; the AES work ran over the course of a week.
I want to put that number in context, because reactions to it split neatly and both sides are missing something.
One camp reads $100,000 as expensive. Compared to a single API call it obviously is. Compared to funding a specialist cryptographer for the two years HAWK had already been under review, it is a rounding error — and the two years of review did not find this.
The other camp reads it as a bargain and concludes original research is now cheap. That overshoots. Anthropic spent an enormous amount of institutional expertise selecting the target, building the harness, and validating the output over hundreds of hours. The $100,000 is the visible portion of a much larger cost.
The defensible reading is narrower and more useful: for a well-specified problem in a formally verifiable domain, with a competent harness and a human willing to push, novel expert-level output now has a price and that price is in the low six figures. It was previously unavailable at any price on that timescale.
The direction of that number over the next eighteen months is the thing to watch, particularly given that OpenAI cut its cheapest tier by 80% this same week. Cryptanalysis is not running on the cheap tier — but everything gets cheaper eventually, and the interesting question is what this costs in 2028.
The Follow-Ons, Including One That Is Not Theoretical
Anthropic reported that after these two results it broadened its search and found more:
- A practical attack on 13-round LEA that, per The Next Web's summary, recovers keys in under an hour on a desktop machine.
- A key-recovery attack against six of Serpent-128's 32 rounds.
- Smaller improvements involving Salsa20, Poseidon and SHA-1.
- A new benchmark, CryptanalysisBench, built with researchers from ETH Zurich, Tel Aviv University and the University of Haifa to measure how well language models analyse cryptographic systems — released so others can evaluate the same capability independently.
The LEA result deserves a second look precisely because the word attached to it is practical rather than theoretical. It still concerns a reduced-round variant rather than a deployed configuration, but it is a different category of finding from a 400-octillion-message thought experiment.
The Next Web also reports that Claude Mythos found 10,000 critical software vulnerabilities in a single month, and frames the move from implementation bugs to algorithmic flaws as a qualitative shift. I would treat that figure as reported rather than independently verified, but the direction is consistent with everything else visible this month.
What This Means If You Run a Business
You are not doing cryptanalysis. Here is what actually transfers.
1. Stop Treating "The Model Says It Can't" as a Finding
When a model tells you a task is not feasible, you have learned something about the model's assessment, not about the task. The most-studied cipher in existence turned out to have something left to find after an articulate refusal.
Practical version: before you abandon an automation because the model would not do it, try a different framing, a different decomposition, and a different model. The multi-agent detail here — one instance dismissed the idea, another exploited it — is a direct argument for not trusting a single response on a consequential question.
2. Invest in the Loop, Not the Model Subscription
The differentiator in this story was a scaffold that let the model form hypotheses, test them, and iterate against real feedback. Not a bigger model. Not a longer prompt.
Most business workflows I see are single-shot: prompt in, output out, human reads it. The step change comes from closing the loop — giving the system a way to check its own work against something real before a human ever sees it. That is also the only reliable route to a high Automation Ratio, because a system that validates its own output is a system whose output needs less correction.
3. Domain Expertise Is Not the Bottleneck You Think
A non-specialist directed frontier-level work in a specialist field. If you have been telling yourself you cannot automate something because nobody on your team is an expert in it, that constraint is weaker than it was.
What you do need is someone who can specify the problem precisely, judge whether an output is any good, and keep going. That is a different skill from domain mastery, and it is far more transferable.
4. Match the Model to the Job, Still
Nothing here argues for running your invoice processing on a research-preview model at $100,000 a go. Cryptanalysis is the archetypal case for maximum capability: an unbounded search over a formally verifiable space where a single success is worth enormous effort.
Almost nothing in your business looks like that. Most of it looks like classification, routing and well-specified transformation, which is exactly what task-model matching says to run on cheap models — and what the four-category framework says frequently needs a rule rather than a model at all.
5. Ask Where Your Cryptography Lives
The one genuinely practical security action, and it comes from the CyberScoop coverage rather than from me: know where cryptography sits inside your systems, what it is connected to, and whether you have a post-quantum readiness plan.
Most small and mid-sized businesses cannot answer the first part of that. You do not need to act on HAWK — nobody deploys HAWK. You do need to know what you are using, because the pattern this week establishes is that algorithms will now be re-examined faster than they were.
The Part That Should Give You Pause
I have been positive about this result, because it is a genuine scientific contribution disclosed responsibly. Two things temper that.
First, Anthropic's own warning: future models may discover flaws with immediate real-world consequences. This time the targets were an undeployed candidate and a reduced-round research variant, and the disclosure ran through NIST. The same capability aimed at deployed cryptography by someone with no interest in disclosure is a different story with no press release.
Second, and more immediate — the model that did this is not one you can buy. Mythos Preview is unreleased, available to a small number of trusted organisations. That asymmetry is now a structural feature of the landscape rather than a temporary state, and it is the same capability-threshold dynamic that has been shaping government access all year.
Set this beside the rest of the month — JADEPUFFER, the supply-chain attacks on developer tooling, two labs disclosing that their own models breached production systems — and the shape is consistent. Offensive and defensive capability are both compounding, and the gap between what these systems can do and what ordinary security assumes they will do keeps widening.
That is not a reason to panic. It is the argument for governance-first architecture stated by events rather than by me.
The Line I Keep Coming Back To
A model, asked to do something hard, explained clearly and persuasively that it could not be done. A human who was not an expert in the field disagreed, built a better loop, and asked again.
Sixty hours and $100,000 later, a scheme that had survived two years of specialist review had lost half its security margin.
Nothing about that requires you to care about lattices. It requires you to notice that the constraint was never the model's capability — it was the structure around it and the willingness to push past a confident no. That is a process ownership story wearing a cryptography costume, and it is available to any business that decides to build the loop rather than buy the tool.
Satya Nadella's point about the learning loop applies exactly here. The organisations that pull ahead are not the ones with the best model access. They are the ones that iterate faster on the structure around it.
Frequently Asked Questions
Is my data at risk because of this?
No. Full AES-128 and every deployed encryption standard are unaffected. HAWK is not used anywhere — it is a candidate under NIST evaluation. The AES finding applies to a seven-round research variant and requires around 400 octillion chosen messages, which is not a practical attack by any definition. If you change nothing about your security posture as a result of this story, you have responded correctly.
Does this mean AI will break encryption soon?
Nobody credible is claiming that, including Anthropic. What it establishes is that AI can now find genuine algorithmic weaknesses that expert humans missed, which is a real capability milestone. Whether that trajectory reaches deployed cryptography, and on what timescale, is unknown — and treat confident predictions in either direction with suspicion.
Should I be worried that AI can do original research now?
Worried is the wrong frame; interested is the right one. This is one result in a formally verifiable domain where correctness can be checked mechanically — which is the easiest possible setting for this kind of contribution. Extrapolating to messier fields where truth is harder to verify is a much bigger leap than the headlines imply.
Can I use a model this way in my own business?
The pattern, yes — the specific model, no, since Mythos Preview is unreleased. But the transferable method is the hypothesis-test-iterate loop rather than single-shot prompting, and that works with models you can access today. Start with one workflow where the system can check its own output against something real before a human sees it.
Why did the AI refuse at first, and does that happen to me?
It assessed the problem as intractable based on the same reasoning any well-read expert would apply — AES is the most-studied cipher there is. It happens constantly in ordinary use, and most people accept it. The lesson is to treat a model's feasibility judgment as one data point rather than a verdict, and to re-ask with a different framing or a different model before giving up.
What is the single most useful takeaway?
That the same model produced a confident refusal and a novel mathematical result, and the only thing that changed between them was the structure around it and a human who kept going. Your results are determined far more by that structure than by which model you subscribe to.
Related Reading
You Don't Have an AI Problem, You Have a Systems Problem — the argument this story demonstrates at the outer edge of technical difficulty.
Bending Spoons: $2.57M Revenue Per Employee — humans on specification and judgment, machines on production, now visible in cryptanalysis.
Stop Chasing the Biggest Model — why the model that does research-grade work is not the model that should run your invoicing.
The Automation Ratio Is the Only Metric That Predicts Survival — and why closing the validation loop is the only reliable way to move it.
The 'AI Business' Advice Is Wrong: Tool Delivery vs Process Ownership — building the loop versus buying the tool, which is the whole distinction here.
The Five Eyes AI Agent Security Guide — governance-first architecture as capability on both sides keeps compounding.
The Government AI Threshold — the access asymmetry that puts the model behind this result out of commercial reach.
About the Author
Hamza Baig is the founder of Hexona Systems, an AI automation agency serving clients across six continents, and the AI Automation Institute, a community of more than 40,000 entrepreneurs building with AI.
He has been featured in the GHL Top 50, Yahoo Finance and Brainz Magazine, and writes regularly on automation architecture, agent governance and the operational realities of AI deployment.
Read more analysis on the Hamza Automates blog, or get in touch to discuss an automation build.
Follow @hamza_automates on Instagram for daily automation breakdowns.
Note: this article summarises published cryptographic research for a general business audience. Technical details are simplified; consult Anthropic's original publication and the underlying disclosures for precise claims.
About
Hamza Baig is the founder of Hexona Systems—an automation agency and softwareplatform that helps thousands of entrepreneurs and business owners implement AI-powered workflows at scale.








