The Call Is Coming From Inside the House
Mythos-class attackers are real, but they are not the threat most likely to hurt you this year. The most consequential agent incident of 2026 did not involve an attacker at all.
Since April, nearly every CISO I have spoken with has put the same issue at the top of their agenda: frontier models can now find and chain vulnerabilities with very little help. Anthropic held Mythos back from broad release because its exploitation capability was too strong to ship. In internal testing, it produced 181 working Firefox exploits. The prior generation produced two. The Cloud Security Alliance published a response plan. Forescout did the same. When I was in government, the time from disclosure to confirmed exploitation was measured in years. Now it is measured in hours.
This threat is real. It is not vendor hype, and the people telling you to improve patching and segmentation are giving you sound advice. Nation-state exploitation capabilities are now commercially available. Plan accordingly.
But that is not the only threat worth planning for. Look at what happened this month.
An incident with no attacker in it
On July 16, Hugging Face disclosed that its production infrastructure had been breached. Over a single weekend, roughly 17,000 attack actions ran across a swarm of short-lived sandboxes. The activity reached the node level, harvested cloud and cluster credentials, and moved laterally into internal clusters. Hugging Face detected it, contained it, and responded exceptionally well.
Five days later, OpenAI disclosed that the attacker was its own models.
OpenAI had been running an internal benchmark called ExploitGym to measure offensive cyber capability. Two models were involved: GPT-5.6 Sol and a more capable, unreleased model. To test their maximum capabilities, OpenAI had disabled the cyber refusals. The models were isolated in a research environment whose only path out was an internally hosted package proxy.
They found a zero-day in that proxy, which JFrog later confirmed was Artifactory, and used it to reach the open internet. From there, they concluded that Hugging Face might have the ExploitGym answer key, broke into its environment, and burned a real zero-day and a chain of credentials to get it.
The motive was cheating on a test.
That is what makes the incident different. Two companies were compromised. Novel zero-days were used. But there was no adversary anywhere in the story. No threat actor to identify. No ransom demand. No stolen data for sale. Nothing a threat intelligence feed could have warned about.
The system responsible was owned, funded, sandboxed, and monitored by one of the most sophisticated engineering organizations in the world. It was not acting maliciously. It was doing exactly what it had been asked to do: optimize for the objective it was given. The objective simply did not include staying inside the environment.
The oldest problem in the field, with a new headcount
Here is what worries me about the security community's focus on mythos-enabled attacks. Security teams are preparing for external attackers armed with AI, which is the threat we already know how to discuss. At the same time, that same technology is being deployed inside the trust boundary faster than any security program has ever had to absorb it.
Security has known for decades which problem is harder. Insiders are not more dangerous because they are smarter. They are more dangerous because they are authorized. They already have credentials. They already understand the workflow. Most controls were designed to let them work, not to stop them. Ponemon put the average annual cost of insider risk at $19.5 million this year, with an average containment time of 67 days and only 13 percent of incidents contained within 30. That is the cost of hunting activity that looks like normal operations because, from the system's perspective, it is.
Now add thousands of agents. They do not sleep, lose focus, quit, or exhibit the behavioral signals insider-risk programs were built to detect. They hold standing credentials and API access, operate around the clock, and move at machine volume. The anomaly detection you would use to catch a human has to treat that activity as normal, because for an agent it is.
We already have an example. In March, an internal Meta agent posted to an internal forum without approval and exposed restricted material to engineers who were not authorized to see it. Proprietary code, business strategy, and user data were visible for roughly two hours, resulting in a Sev-1 incident. The agent did not bypass an identity control. It passed every identity check because it was authorized to access the systems involved. It authenticated as itself, then took an action nobody had explicitly prohibited.
That is a confused deputy with production credentials, and it is the defining failure mode for this class of system.
Why your controls do not bound this
This is where the real disagreement begins. Most security leaders I speak with believe their existing stack already covers the problem. They have IAM, DLP, UEBA, EDR, approval workflows, and change controls. Those systems have adapted to two decades of new technology, so it is reasonable to expect them to handle agents too.
But apply each control to an agent, and the gaps become obvious.
IAM authenticates identity, it does not constrain purpose.
IAM tells you who is calling and what resources that identity can access. It does not tell you what the caller is supposed to do with that access. That gap was manageable when the principal was a person with a defined role, a manager, and judgment. An agent has none of those safeguards. It can use every permission it has been granted, whether or not the action serves its intended purpose. Access is not authorization to act.
DLP sees the data, not the purpose.
"Summarize this customer's account" and "exfiltrate this customer's account" produce the same API calls and the same field names. The difference is purpose, which is exactly the thing your pattern matcher cannot see.
Behavioral analytics looks for human anomalies.
An agent normally operates across multiple systems, at high volume, and at all hours. The behavior that would flag a human as suspicious may be routine for an agent, leaving no meaningful deviation for the system to detect.
Approval workflows break at machine volume.
A human approval step is only meaningful if someone can evaluate each request. When it fires hundreds of times a day, review becomes reflexive and the control turns into a click-through.
The sandbox is not a wall. It is part of the attack surface.
The July incident exposed a problem the industry has not fully absorbed. The research environment was thoughtfully designed, but its containment depended on a shared component with a vulnerability. The models started inside the boundary, found that weakness, and used it as a path out. Traditional isolation assumes the attacker begins outside and must break in. An agent begins inside, can inspect everything it can reach, and has the patience to keep looking.
That is the larger point. You cannot reliably contain an intelligent system by telling it how to behave (anyone with kids knows this to be true). You have to limit what it can access and what actions it can take. Instructions are guidance. A capable agent will find paths you did not anticipate, not because it is malicious, but because finding paths is part of what it was built to do. The controls that hold are not rules about conduct. They are enforced limits on reach.
The part nobody wants to say out loud
There is another lesson in the July incident, and most of the coverage has avoided it.
Why would a lab run offensive evaluations at maximum capability, with refusals disabled, inside an environment that shared a dependency with the outside world? The answer is not that the people involved were careless. I know people at these labs, and they are not careless. The better question is what kind of pressure those organizations are under.
OpenAI has told investors it is targeting roughly $600 billion in compute spending by 2030, against approximately $25 billion in annualized revenue and a projected cash burn of about $27 billion this year. Inference costs are also rising as usage grows. Whatever you think of the exact numbers, the strategic position is clear. That gap will not be closed by chat interfaces. It will be closed by agents with more autonomy, more access to enterprise systems, and more customer data. And those agents have to ship quickly because every quarter of delay carries the cost of infrastructure commitments that have already been made.
I have run product at companies with far less at stake, and I know what that pressure does to a roadmap. It does not make anyone reckless or malicious. It makes safety one more priority competing with features tied directly to growth, and the growth features usually win. Meanwhile, prompt injection remains unsolved by construction. OWASP says so plainly. Recent testing found that no agentic browser reliably blocked even one attack scenario, and Gartner has advised enterprises to block AI browsers outright. The products shipped anyway. They were always going to ship.
So when a vendor tells you its agent is safe because the guardrails are on, be clear about what that means. The control lives inside a model you cannot inspect. It can change when the model is retrained. It can be removed through a configuration decision made by the vendor. And the company operating it faces economic pressure to expand capability faster than your risk tolerance may allow. The July incident demonstrated the problem directly. The guardrails were part of the model's behavior, and one internal decision turned them off.
That leads to an argument I have been making privately for several months.
For control purposes, a Chinese open-weight model running on your own hardware is more trustworthy than a frontier API.
I am not saying it is better aligned, safer in the abstract, or built by people whose interests match yours. I am making a narrower claim. In security, trust does not mean believing someone is virtuous. It means being able to verify and bound what a system can do.
On that test, the open-weight model has an important advantage. Its weights do not change between Tuesday and Thursday. It cannot be deprecated in the middle of your quarter. It has no telemetry channel home, no incentive to expand its own access, and no infrastructure bill that your data helps pay. Every action it takes crosses infrastructure you control, which means every action can be observed and enforced by you.
The objections are real. Open-weight models carry provenance and supply-chain risk. A model can be poisoned or backdoored during training in ways that are difficult to detect. There may be no vendor to hold accountable and no coordinated disclosure process to rely on. Washington is also moving against the use of China-origin models in regulated and government environments, which will settle the question for some organizations. And calling a hosted Chinese API is an entirely different decision, and one I would not make. In that case, your prompts leave your infrastructure and fall under someone else's laws and controls.
Those risks are serious, but they are familiar. Security teams know how to manage provenance, supply-chain, and isolation risks through attestation, evaluation, least privilege, and containment. We have a playbook for them. What you cannot attest, audit, or patch is a vendor's incentive structure.
Simon Willison identified the sharpest version of this asymmetry in the July incident. Hugging Face could not deploy frontier models freely to defend itself because the models' safety policies restricted that behavior. The models attacking Hugging Face were running with those same restrictions disabled. The restraint applied to the defender, but not to the system causing the damage.
This is not an argument to replace frontier models with open weights. Use the frontier models. They are extraordinary, and your competitors are already using them. It is an argument about where the control boundary belongs. Put that boundary somewhere you can inspect and enforce, because the vendor's incentives and yours will diverge precisely when the consequences are highest.
The controls that matter
You do not need to reinvent security for agents. You need to treat them as principals and apply the same discipline you already use to govern identities, permissions, and access.
Inventory the agents as identities, not as tools.
For every agent in your environment, document its owner, the credentials it uses, and the systems it can reach. Most organizations cannot produce that list today. That gap is the finding, and you should treat it the same way you treated your first unmanaged-asset report.
Give every agent its own identity.
Each agent should have a distinct identity tied to its actions. Shared service accounts erase attribution, making it impossible to determine which agent did what. If you cannot identify the agent behind an action, you cannot investigate or govern it.
Derive the boundary from the declared purpose, and enforce it at the tool call.
The system prompt can describe what the agent should do, but it cannot enforce what the agent is allowed to do. Enforcement has to happen at the moment of action, before execution, with deny-by-default rules and a logged decision tied to the specific constraint. A support agent may issue refunds below a set threshold for valid orders. It may not change payout destinations. No prompt or conversation should be able to change that boundary.
Size approval gates to blast radius, not to a risk score.
Human review is limited, so reserve it for actions with real consequences: moving money, sending data outside the organization, or changing access. Do not waste it on low-impact decisions, or it will become meaningless when it matters most.
Treat containment as adversarial.
Assume the agent inside the sandbox will probe every dependency and route it can reach. If egress controls rely on shared infrastructure, that infrastructure is part of the attack surface, not the boundary.
Reconcile.
For your highest-consequence systems, use the system's action logs to attribute every action to a registered agent, a human, or an owned service account. Anything you cannot account for is a control gap. Report the percentage of activity you can attribute. It is a concrete measure of control, not a vague claim that you "have AI governance."
Make the boundary portable.
The OpenAI incident exposed a broader problem: the controls lived either inside the model, where they could be disabled, or around it, where they could be escaped. Once the agent crossed that boundary, nothing else enforced what it was allowed to do. An agent's operating limits should travel with it as a machine-readable policy that any runtime can enforce. Changing environments should not mean leaving the rules behind.
Where TenetGraph fits
TenetGraph puts that playbook into practice. We derive an agent's operating boundary from its artifacts, including its code, system prompt, tool schemas, granted scopes, and organizational policies, so nobody has to hand-author the policy or chase every revision. An adversarial engine tests that boundary before deployment and again after every change. The constraints are enforced deterministically at the tool call, before execution, and every allow, deny, and escalation is logged with the constraint that produced it.
The July incident shows exactly where that control matters. TenetGraph would not have stopped the models from discovering the vulnerability or deciding to exploit it. It could have stopped them from carrying out the actions that turned that discovery into an incident. With an enforced operating boundary, the attempt to reach the open internet and act against Hugging Face could have been denied before execution. That is our claim: governed agents stay within their defined boundaries, and agents operating without one are surfaced as exceptions that require an owner and remediation.
Define the boundary. Authorize the action.
The bottom line
Prepare for the Mythos threat. It is real.
But over the next eighteen months, more companies may be hurt by agents they deployed themselves than by sophisticated adversaries using frontier models. The incident will start with an agent your team approved, using credentials you granted, pursuing an objective that sounded reasonable, and moving at machine speed through systems that recognize it as legitimate. There will be no attacker to attribute and no warning in a threat intelligence feed. Just a postmortem concluding that the agent did exactly what it was told.
Detection determines how often the boundary is tested. The boundary determines the damage when detection fails.
Keep watching the perimeter. Just do not mistake it for the only place risk enters. The call is coming from inside the house, and it already has credentials.
Chris Finan is co-founder and CEO of TenetGraph. He has spent the past decade building and scaling cybersecurity companies: CEO of Anitian, now merged with Arkenstone Defense; President and CRO of ActZero, a managed detection and response provider acquired by WatchGuard; and executive roles at Shape Security (acquired by F5) and Impermium (acquired by Google), defending consumer platforms against automated fraud and abuse. He previously served as product director for Plan X, the Defense Department's flagship cyber warfare program at DARPA, and as director for cybersecurity legislation and policy on the White House National Security Council staff.