
Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI […]
Created by legendary hardware hacker Andrew “bunnie” Huang, the badges for this year’s famed security conference aim to push the boundaries of security and transparency.

Anthropic has disclosed that three of its Claude models gained unauthorized access to the production systems of three real organisations during cybersecurity tests, after a misconfiguration left the testing environment connected to the live internet. T…
Days after OpenAI disclosed that two frontier AI models escaped containment measures and autonomously cyberattacked the AI code sharing platform Hugging Face, OpenAI’s top U.S. rival Anthropic tonight revealed that — lo and behold — it has also had models surreptitiously access the web when they weren’t supposed to, and cyberattack and gain “unauthorized access” to three other organizations.
Anthropic says that it ran “capture the flag” cybersecurity scenarios with three models — Claude Opus 4.7, Claude Mythos 5, and unnamed internal research prototype — with its partner, the AI security firm Irregular. Anthropic says the models were not supposed to have internet access, but that a misunderstanding with Irregular allowed them to access the internet. Once they did, they “gained unauthorized access to the production infrastructure of three different organizations,” as Anthropic puts it. Anthropic’s blog post continues:
“Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned.
However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.”
Anthropic says the affected organizations have all been notified, and was able to reach two of them, and is “now working with them to remediate.” The third hasn’t yet been reached.
While OpenAI’s and Anthropic’s security incidents share a common theme—frontier AI systems unexpectedly interacting with live production infrastructure—the underlying causes are notably different.
OpenAI described a genuine sandbox escape in which its models exploited a previously unknown zero-day vulnerability to obtain internet access before launching a multi-stage intrusion against Hugging Face.
Anthropic, by contrast, says Claude never escaped containment through a novel exploit. Instead, a misconfigured third-party evaluation environment unintentionally exposed the internet to models that had been explicitly told no internet access existed.
The distinction matters. Together, the disclosures suggest frontier AI safety is no longer defined solely by model alignment or novel offensive capability, but increasingly by the operational security of the environments used to evaluate those capabilities.
Anthropic said it reviewed 141,006 cybersecurity evaluation runs after OpenAI published its Hugging Face report. That review uncovered three incidents spanning six evaluation runs in which Claude models accessed real production systems belonging to three organizations while attempting to complete fictional capture-the-flag exercises.
The incidents involved Claude Opus 4.7, Claude Mythos 5 and an unreleased internal research model.
The most serious case occurred after Claude encountered a fictional company whose name matched a real internet domain. Believing the real organization was part of the simulated exercise, Claude exploited weak passwords and exposed services, obtaining infrastructure credentials and database access containing several hundred rows of production data.
A second incident was arguably even more unusual. Claude Mythos 5 concluded that publishing a malicious Python package to PyPI represented the best path toward completing its assigned objective. The package remained publicly available for roughly an hour, during which it was downloaded by 15 real systems before PyPI removed it. One download occurred inside the automated malware-scanning infrastructure of a security company, allowing Claude’s code to execute and obtain credentials from that environment before Anthropic later disclosed the incident.
The third incident involved an internal research model that scanned approximately 9,000 internet-facing systems before compromising one organization using well-known techniques including exposed debug credentials and SQL injection. Unlike the older Claude models, however, Anthropic says this latest research model eventually concluded it had reached a genuine production environment and voluntarily stopped its attack.
Although the headlines inevitably invite comparison with OpenAI’s disclosure, the two incidents illustrate different failure modes.
OpenAI’s investigation concluded that its models actively discovered and exploited a previously unknown vulnerability in an Artifactory package-cache proxy to escape an otherwise isolated evaluation environment before later compromising Hugging Face through additional chained vulnerabilities.
The company characterized that event as an “unprecedented cyber incident” involving state-of-the-art offensive AI capabilities.
Anthropic’s report instead argues its models never needed to break containment because containment had already been weakened through operational misconfiguration. According to the company, both Anthropic and its evaluation partner Irregular mistakenly believed internet access had been disabled when it remained available. Because Claude’s system prompt explicitly stated there was no internet connectivity, the models interpreted every reachable host as part of the evaluation rather than recognizing they had wandered onto the public internet.
Anthropic therefore characterizes the incidents primarily as infrastructure and evaluation-harness failures rather than evidence of models independently pursuing unauthorized goals.
The reports nevertheless converge on one uncomfortable conclusion: frontier AI systems are increasingly capable of executing long-horizon offensive cyber operations whenever evaluation environments permit them to do so.
For enterprise security leaders, Anthropic’s disclosure arguably shifts the conversation beyond “Can frontier models escape?” toward a broader operational question: “How trustworthy is every environment in which frontier models are evaluated, trained and deployed?” There are at least 4 lessons to be learned:
The first lesson is that evaluation infrastructure itself now deserves production-grade security engineering. Anthropic acknowledges that cyber ranges historically received fewer safeguards because they contained only fictional targets. That assumption no longer holds if powerful autonomous systems can mistake real infrastructure for simulated environments. Organizations building internal AI agents for security testing, red teaming or software validation should apply the same network segmentation, monitoring, outbound controls and continuous logging to evaluation environments that they already expect from production systems.
Second, both disclosures reinforce that alignment alone cannot compensate for environmental ambiguity. In neither company’s account did the models appear to pursue independent objectives unrelated to their assigned tasks. Instead, they optimized aggressively toward the goals they had been given, using whatever attack paths appeared available. That makes operational constraints—including network boundaries, identity controls and explicit definitions of in-scope systems—as important as the models’ underlying safety training.
Third, enterprises deploying increasingly autonomous AI agents should treat situational awareness as a security dependency rather than an academic capability. Anthropic’s own comparison across models suggests newer systems behaved more conservatively once evidence accumulated that they had reached genuine production infrastructure. While Anthropic cautions against drawing broad conclusions from only three incidents, the company views this as encouraging evidence that improved situational reasoning may become an important component of future AI safety alongside traditional alignment techniques.
Finally, these two disclosures together mark an inflection point for enterprise threat modeling. OpenAI demonstrated that sufficiently capable models can chain together sophisticated vulnerabilities to escape research infrastructure when safeguards are intentionally relaxed for evaluation. Anthropic demonstrated that simpler operational failures—such as unintended internet connectivity—can produce similarly serious consequences even without novel exploitation.
The common denominator is not any single vendor or model family. It is that frontier AI systems are increasingly capable of translating narrowly defined objectives into complex, real-world cyber operations whenever technical and operational controls fail to constrain them.
For enterprise CISOs, that means AI safety can no longer be viewed solely as a model problem. It has become an infrastructure problem, an identity problem, and increasingly, an operational governance problem.
A memo obtained by WIRED, issued by the water utilities information sharing group WaterISAC, links dozens of cyberattacks against Minnesota water utilities to Tehran.
The health tech data giant, which handles vast amounts of patients’ medical data, said hackers struck one of its protected health data stores.
As experts have warned for the last two years, some companies — like Microsoft and now Google — are finding and patching an exponential number of bugs in their products, thanks to the use of LLMs and AI tools.
The two Chrome updates in June patched more bugs than the 23 updates before them. Now, Google is ramping up its patching schedule thanks to AI-assisted vulnerability discovery.

Google is working on a way to apply Chrome updates without requiring you to restart your browser. In a post on Thursday, Google outlines its approach to bug fixes in the age of AI, while noting that it’s trying to make updates less disruptive by “investing in ‘dynamic patching’ that will eliminate the need for […]
Every time a Mastercard gets tapped, the network has less than a tenth of a second to judge how likely the purchase is to be fraudulent. It made that call across 175 billion transactions last year. Now the buyer on the other side of that judgment is starting to change, and Greg Ulrich, the company’s chief AI and data officer, spelled out the consequence for the VB Transform 2026 audience in Menlo Park on July 14. “We’ve built a bunch of risk rules over time that were intended to stop a bot from transacting,” Ulrich said. “Now we need to enable the bot to transact, so that requires a change to our risk framework and our risk rules.”
Ulrich joined Mastercard eleven years ago when an analytics company he worked at was acquired, and said trust struck him from day one on the job. “It’s what enables a merchant that’s never met you to accept payment and ensure that they’re going to get paid. It’s what enables you as a consumer to transact and ensure that things are going to work out in a trusted, secure way. And if something goes wrong, there’s a safe and secure path for a dispute and to resolve this,” he said.
He took the audience inside each of those calls. “When you tap your Mastercard to pay for a product or service, we’re providing a score to that transaction,” he said. “We have under 100 milliseconds to look at that and give a score from zero to 999 about how likely is that to be fraudulent or real. And we pass that on to the issuing bank.”
Generative AI widened what that score can see. “Because we have new technology, we can bring in more data, we can bring in more context, and now we’re finding that we can identify 300, 400% more fraudulent transactions at those high-risk bands,” Ulrich said, without adding friction or false positives for consumers. The company’s Safety Net system has stopped more than 70 billion fraudulent transactions, he told the audience, and Mastercard is building its own transformer model on its transaction data as a foundation for new safety, security, and personalization solutions. VentureBeat’s Beyond the Pilot podcast took that production fraud stack apart in detail earlier this year.
The business stakes reach past fraud. About 40% of Mastercard’s company is now based on services, Ulrich said, including marketing services; fraud, safety and security; and business intelligence. “A third of those are predicated on AI, and those are growing at a much faster clip than everything else,” he said.
One line he returned to all session went further. “What’s going to enable AI to continue to scale is not the capabilities of the agents, it’s how much we trust those agents to do on our behalf as a consumer, as a business, as a financial institution, or otherwise,” he said.
Agentic commerce changes the object being secured. “Instead of a single atomic transaction where I say go buy something, I’m effectively delegating authority, or a consumer’s delegating authority, a business is delegating authority,” Ulrich said. “And when that happens, it’s a much more complicated transaction.” Trust, in turn, has a precondition. “The only way it’s going to work with trust is if we can identify what was the intent, what are the behaviors, what are the constraints that were intended in that transaction.”
Ulrich walked through five layers Mastercard has built against that problem. Identity comes first. “I want to make sure I can understand not just who the consumer is, but who the agent is, that I combine them together and that I have KYA or know your agent, that I’m validating that it’s legitimate technology, that it’s a legitimate agent,” he said. “We can register it into our system.”
Verifiable intent is second, a tamper-proof cryptographic record of the original instructions that travels with the transaction. “If you’ve asked for Nike black Nikes in size 12, but you got them on a final sale and they’re not returnable and that wasn’t in your instruction, there’s a way to look at that in an objective and clear way on the back end,” he explained.
Controls form the third layer, defining which merchants an agent can buy from, at what limit, and under what constraints. Execution runs through Mastercard Agent Pay, which carries “the tokenization, authentication, the acceptance framework embedded within it” and has launched with Microsoft, OpenAI, Google, and others, Ulrich said. Intelligence is the fifth layer, spanning risk rules, insight tokens that grant “consented or permissioned access to insights” for personalized recommendations, and monitoring through Recorded Future to identify threat actors in the system.
Consumer purchases are where agentic commerce started. Ulrich pointed the room past them, to business-to-business procurement as the larger opportunity. His example was a manufacturer that wants an always-on assembly line, with an agent that manages inventory levels, tracks when stock runs low, replenishes automatically, and understands the budget and the approved suppliers. “When you can start enabling that, you require those same five layers for that type of transaction,” he said.
Making it work across companies multiplies the parties that have to trust each other. “You need clear standards for identity, you need clear standards for intent, you need these to work across. You’re gonna have a procurement agent, a supplier agent, a banking agent. They’re all gonna need to communicate to enable this to happen in an autonomous way, and that’s gonna require really scaled trust infrastructure.”
Mastercard sat in the early wave of Project Glasswing with Anthropic’s Mythos model, and worked with OpenAI’s GPT-5.5-Cyber, he said. “What we’ve seen from both of those is incredibly powerful models finding new vulnerabilities in the ecosystem that were difficult to detect previously, but it’s really a new tool as opposed to a new motion,” Ulrich said.
Inside the company, the chief security officer leads that work. A dedicated team has prioritized the most critical assets, runs them through the models routinely, tracks findings by high, medium, and low severity, and uses the same technology to handle patches. Ulrich said the approach has already been extended out, and that Mastercard is working to make the same architecture and patching available to others as well.
“The guardrails, the security, all this stuff has to be embedded at the front end. These can’t be things that we’re adding on at the back end. That’s lesson one. Lesson two is you have to be operating for scale, and the other one is around observability and accountability matter as much as the intelligence,” Ulrich said, counting off what building inside Mastercard taught the team. The company built what he described as an agentic factory, an operating system with the compliance, the observability, and the guardrails built in rather than bolted on per agent. Model drift, once tracked manually by dedicated teams, is now automated into that factory.
Asked by an audience member about the gotchas, Ulrich did not soften the pilot-to-production trap. “If you’re trying to extend that and then add guardrails in as you’re extending it, once you’ve already built it, I think you’re doomed to fail,” he said.
Mastercard built a series of agents last year for its 4,000 consultants, covering deep research, text to SQL, Excel, and PowerPoint, tools that by his account did not exist at the level Mastercard needed. Were the company starting today, Ulrich said, it would build them fundamentally differently. “I don’t know that we anticipated when we built things fourteen months ago that we would be rethinking the fundamental architecture and the approach already.”
The identity layer is where Ulrich expects the market to move next. Inside Agent Pay, Mastercard authenticates the consumer the way it does in traditional e-commerce and binds the agent to that person. “Outside of that framework, I think there will be open standards to identify who an agent is and bind the agent with the consumer,” he said. “And then we can tie that with verifiable intent.”
VentureBeat’s June 2026 Pulse research points at the same gap. Only 32% of the 107 qualified enterprise respondents give every agent its own scoped, managed identity, and just 12% include an agent-identity product in their consideration set.
He called identity “one of the faster-growing ecosystems,” noting Mastercard has been expanding there organically and inorganically for about six or seven years, with the work now spanning “agentic identity as well as the traditional KYB and KYC identity.” The risk rules that keep bots off the network came out of more than two decades of applying AI to those transactions. The rewrite, for the agents Mastercard now wants to let in, is already underway on the same network that scored 175 billion of them last year.