TL;DR
White House AI adviser David Sacks says Anthropic refused to fix a jailbreak that could restore high-end cyber capability in Fable. Anthropic says the flaw was minor, already known and not unique to its model. The dispute remains hard to verify because the technical evidence, test method and reported outside partner have not been made public.
White House AI adviser David Sacks has publicly defended a U.S. restriction on Anthropic’s Fable models, saying Anthropic refused to fix a jailbreak that could restore Mythos-level cyber capability; Anthropic disputes the account, saying the flaw was minor and not specific to its systems. The fight matters because the technical evidence, the test method and the outside party said to have flagged the issue have not been made public.
Sacks, co-chair of the President’s Council of Advisors on Science and Technology, wrote on X on June 13 that a "highly credible trusted partner" found a bypass of Fable’s guardrails. He said the administration asked Anthropic CEO Dario Amodei to fix the issue or pull the model, and that an export control was issued after he refused. Those points remain Sacks’s account, not an independently verified public record.
Anthropic’s June 12 blog said the government provided no specific technical detail, that the demonstration showed only a few minor, already known flaws, and that similar outputs could be reproduced on other public models, including GPT-5.5. Anthropic argued that a narrow potential jailbreak did not justify recalling a model used by hundreds of millions of people.
Semafor reporting cited in the source material and carried by Fortune and others identified Amazon as the possible trusted partner, with CEO Andy Jassy reportedly in contact with the administration. Amazon has not confirmed that role or the details of any test. The unresolved point is the severity gap: Sacks describes restored cyberweapon operability, while Anthropic describes a routine safety defect.
The Safety Card, Played From Every Side
● ContestedA White House adviser says Anthropic refused to fix a cyberweapon jailbreak and got banned for it. Anthropic says the flaw is trivial. Almost every fact that would settle it is non-public — and “safety” is now the card every side is playing.
Both are claims, not findings. They don’t disagree on tone — they disagree on what the bypass actually is.
- A “highly credible trusted partner” found a jailbreak of Fable’s guardrails.
- The admin asked Amodei to fix it or pull the model. He refused.
- So the export control was issued — “reluctantly.”
- It restores operability of a cyberweapon; calling that “not serious” is indefensible.
- The government gave no specific technical detail.
- The demo found a few minor, already-known flaws.
- Other public models (incl. GPT-5.5) do the same without a bypass.
- A “narrow potential jailbreak” shouldn’t recall a model used by hundreds of millions.
Per reporting by Semafor (carried by Fortune and others), the entity that flagged the jailbreak was Amazon — with CEO Andy Jassy reportedly in contact with the administration. Amazon hasn’t confirmed specifics. Flagging a real risk is what a good partner does — but Amazon wears three hats at once, and none of them is neutral.
Each actor’s safety claim points toward its own advantage.
The entire evidentiary record is a matter of trusting parties who each have a reason to shade it.
A transparent, technically grounded, independently reviewable process — which is, notably, exactly what Anthropic says it wants, and exactly what would also constrain Anthropic. The reason to demand it isn’t loyalty to anyone; it’s that the alternative is decisions made on secret evidence and adjudicated in dueling press statements.
Independent commentary, produced with AI assistance under human editorial oversight; the views are the author’s own and may change. This is analysis and opinion, not investment, financial, legal, or technical advice, and it concerns an actively developing situation in which key facts are disputed and non-public. Claims attributed to David Sacks reflect his June 13, 2026 statement on X; claims attributed to Anthropic reflect its published statements; reporting on Amazon’s role reflects accounts published by Semafor and others — all read as of June 15, 2026, and presented as the claims of those parties, not as established fact. Characterizations are the author’s interpretation, offered in good faith and open to rebuttal. References to specific people, companies, and government actions are factual and analytical, not partisan, and imply no affiliation or endorsement.
Opaque Tests Shape AI Policy
The case is a test of how government officials, AI developers and major technology partners handle claims about dangerous model capability when the proof is sensitive or held privately. A real bypass that restores high-end cyber capability would give regulators a strong public-safety reason to move fast; a weak or ordinary jailbreak would raise concerns about secret evidence being used to sideline a commercial model.
It also reverses a familiar role in AI safety debates. Anthropic has often argued that highly capable systems can pose cyber risks and need firm rules, but in this dispute it says the government is overstating the danger. Sacks is using the safety rationale to defend the restriction, while a reported Amazon role would add business rivalry and cloud dependency to a dispute already short on public evidence.

Ai Engineering Made Practical: Build Reliable Ai Systems With Retrieval, Tools, Evaluation, Monitoring, And Safety—So Teams Ship Faster With Less Risk
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Anthropic’s Safety Argument Turns Back
Fable is described in the source material as Anthropic’s guarded version of Mythos, a model Sacks says Anthropic itself had characterized as close to a cyberweapon. Under Sacks’s account, the policy concern is not ordinary bad output; it is whether a bypass strips away the guardrails that separate Fable from Mythos-class capability.
Anthropic’s response shifts the argument from model capability to test quality. The company says officials did not provide the technical detail needed to verify the claim, and that the behavior shown in the demo appears across other public systems without a special Fable bypass.
The source article labels the episode contested rather than settled. It says the public record lacks a named partner, published methodology, CVE-style disclosure or independent review that would let outside experts compare the two accounts.
“Sacks said a "highly credible trusted partner" found a jailbreak of Fable’s guardrails.”
— David Sacks, according to his June 13 statement on X
Missing Evidence Keeps Severity Open
It is not yet clear what prompt, exploit chain, model version or evaluation standard produced the claimed jailbreak. It is also unclear whether the government restriction covers one model, a model family, export access, customer access or another operating limit.
The reported Amazon connection remains unconfirmed by Amazon in the source material. Amazon’s relationship with Anthropic also cuts across several roles: investor, cloud provider and model competitor. That overlap does not prove improper conduct, but it makes public review of the safety claim more pressing.
There is no public basis yet to decide whether Sacks’s account, Anthropic’s account or some mixed version is closest to the technical facts. The key question is whether the undisclosed test shows a unique dangerous bypass or a familiar model-safety failure.
Patch Timing May Signal Severity
The next marker is whether the restriction is lifted quickly after a quiet model change or remains in place. A fast resolution after a patch would weaken Anthropic’s claim that the flaw was trivial. A longer standoff would put more pressure on the government to publish a reviewable basis for the action.
Independent technical review would settle more than dueling statements. The test to watch is whether officials, Anthropic or the reported partner release enough method detail for outside experts to reproduce the finding without exposing working cyber abuse instructions.
Key Questions
What did David Sacks say about Fable?
Sacks said a trusted partner found a jailbreak that could defeat Fable’s guardrails and restore Mythos-level cyber capability. He said Anthropic was asked to fix or pull the model and refused.
What does Anthropic dispute?
Anthropic says the government did not provide specific technical detail, that the flaws shown were minor and already known, and that similar behavior can be reproduced on other public models.
Has Amazon’s role been confirmed?
No. Semafor reporting cited in the source material identified Amazon as the possible trusted partner, but Amazon has not confirmed the details described there.
Why is this dispute hard to verify?
The prompt, test method, model version, partner identity and independent technical review have not been released publicly. Readers are left with conflicting statements from parties with policy or business interests in the outcome.
What should readers watch next?
Watch whether the restriction is lifted after a patch, whether the standoff continues, and whether any party releases a reviewable technical account of the alleged jailbreak.
Source: Thorsten Meyer AI