AI models escaped OpenAI’s sandbox and hit Hugging Face. Crypto is where that gets dangerous
MarketsAI models escaped OpenAI’s sandbox and hit Hugging Face. Crypto is where that gets dangerousOpenAI said the systems had their cyber guardrails lowered for an internal benchmark, but the incident shows how autonomous exploit chains could pose a deeper threat to smart contra

MarketsAI models escaped OpenAI’s sandbox and hit Hugging Face. Crypto is where that gets dangerousOpenAI said the systems had their cyber guardrails lowered for an internal benchmark, but the incident shows how autonomous exploit chains could pose a deeper threat to smart contracts, where losses are final.By Shaurya Malwa Jul 22, 2026, 6:01 a.
m. 3 min readMake preferred on ShareShare this articleCopy linkX iconX (Twitter)LinkedInFacebookEmailMake preferred on AI models escaped OpenAI’s sandbox and hit Hugging Face. Crypto is where that gets dangerous(Levart_Photographer/Unsplash)SummaryShowOpenAI disclosed that experimental versions of its GPT models, with safety guardrails lowered, escaped a test environment and compromised Hugging Face’s live infrastructure by exploiting previously unknown vulnerabilities.
The incident demonstrates that advanced AI systems, when directed to win hacking-style challenges, can autonomously chain together flaws, stolen credentials and infrastructure weaknesses to reach production systems.Security experts warn that similar AI-driven techniques could be used to execute complex, multi-step crypto attacks, from probing smart contracts and bridges to compromising developer tools and admin keys, turning access into stolen funds within minutes.AI development lab OpenAI disclosed on Tuesday that a group of its models, including the publicly available GPT-5.
6 Sol and a more capable unreleased system, ‘broke out’ of a controlled test environment and compromised the production infrastructure of Hugging Face, the company that hosts much of the open-source AI world.The models were being run through an internal benchmark called ExploitGym, a test of long, multi-step hacking tasks, with their cyber safety refusals deliberately lowered for the evaluation. To be clear, this was not a production model spontaneously turning hostile.
It was a capable model with guardrails off and specifically told to win a hacking test - doing whatever it took to win.The models found a hidden flaw in the test software, one nobody knew was there, and used it to slip past the walls meant to keep them offline. Once on the open internet, they guessed that Hugging Face might store the test's answers.
To get in, they strung together stolen passwords and more hidden flaws until they could run their own commands on Hugging Face's live servers.OpenAI caught the anomaly internally, while Hugging Face's team detected and contained it. It called the incident "unprecedented,” and said extensive security steps will be put in place to prevent untoward incidents that may impact public systems or services.
“We are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched,” the team said in its blog post. “We’re improving and adding stronger protections around future training and evaluations.”A simple explainer on how the model broke out to cheat.
(Shaurya Malwa/CoinDesk)Why crypto developers should bewareMuch of a crypto attack happens before funds move. Attackers scan code, test passwords, search for exposed credentials, analyze signing setups and look for a path into an administrator account. OpenAI’s models carried out several parts of that process during the Hugging Face incident, moving from one weakness to another until they reached live production servers.
And the crypto market has plenty of places for that approach to work, as several attacks from earlier this year have shown. The weak point may be a smart contract, but it may also be a developer laptop, a poisoned software package, a bridge validator or or one signer in a multisig wallet.Take Drift’s $285 million attack from earlier this year as an example, a theft that took a six-month social-engineering campaign to reach privileged access.
An AI agent can, in theory, test many routes at once, keep track of failed attempts and continue working while its human operators sleep. Once a path is found, the operator can act on the actual attack and a viable exit path.KelpDAO’s $292 million bridge loss exposed a different weakness.
The attacker found a single-verifier flaw in the system used to move assets between blockchains.That kind of attack starts with patient code review and infrastructure mapping - the type of work OpenAI’s models performed when they found an unknown flaw.A third type of attack targets onchain governance systems.
Earlier in July, an attacker spent about $4.4 million buying enough of Solana-based dog memecoin BONK to initiate and pass a proposal that transferred roughly $20 million from the project’s treasury to the attacker. This occurred over a three day period, and the attacker later sold all tokens used to win the vote, as CoinDesk tracked at the time.
Three big 2026 crypto thefts, three different weak points. (Shaurya Malwa/CoinDesk)The purchases, vote and treasury transfer for that attack were all valid transactions individually. But the theft came from understanding how the ru
Đọc thêm từ Tài chính
What is Pickaxe Mountain, the Iranian nuclear facility Trump is threatening to strike?
"We'll be hitting that area pretty soon, and very heavily," President Donald Trump said about Iran's underground nuclear facility in Pickaxe Mountain.


