The Agents You Didn't Build

Over one weekend in July, a set of OpenAI's own models slipped the sandbox they were sealed inside, reached the open internet, and broke into another company's servers. To the models, this was a game. 

The goal was to solve a cybersecurity benchmark. Once the models got onto the open internet (which they weren't supposed to reach), they found vulnerabilities in Hugging Face's infrastructure that allowed them to obtain the cybersecurity test solutions. Plainly: they broke in to steal the answers.

Hugging Face detected and contained it, but only after the agent had carried out more than seventeen thousand attacker actions across its systems. And OpenAI didn’t connect its own logs to the breach until about a week later- only after Hugging Face publicly disclosed it. 

In April 2026, Anthropic researchers put an early version of Claude Mythos in an isolated sandbox and instructed it, as part of a safety test, to try to escape. It built a multi-step exploit, broke out, and gained internet access. They only found out because they had asked the model to email the researcher if it was successful (which it did).

But then the model went even further. Without being asked, it posted descriptions of its own exploit on several public websites. Anthropic characterized it not as a malfunction but as an agentic capability operating without adequate goal constraints, they called the model "reckless," and decided not to release it publicly.

So the organizations that build these models, that employ the people who understand them better than anyone alive, that have every resource and incentive to keep their own systems contained, could not reliably keep their own agents inside the box. 

If the frontier labs cannot fully predict or contain what their agents will do, what exactly is the basis for the confidence with which the rest of us are deploying them?

Nobody knows how to do this yet

I have spent the better part of two years working on enterprise agent platforms, on the architecture, the governance, the registries, the reference models. All this has taught me how little most people know.

This is a revolutionary kind of technology. We are deploying systems that take actions on our behalf, adapt their own behavior, and pursue goals in ways their own creators can’t fully anticipate- all at an unprecedented speed in enterprise software. Think about how much AI has changed in the last three years, then try to say with confidence what it will do in the next six months. I can’t. I don’t believe anyone can, and I’ve learned to be suspicious of the certainty of the people who claim otherwise, because that certainty is usually attached to a sale.

Every other week I see another company promising “enterprise agent governance.” I don’t think they’re lying but I think the problem itself hasn’t stabilized yet. 

The exposure you didn’t choose

When executives worry about agent governance, they picture the agents their company builds. But those are now just a small fraction of the agents your company runs. In 2026, every major enterprise platform turned itself into an agent platform. Salesforce, Microsoft, ServiceNow, SAP, Workday, Oracle, Databricks, Snowflake. The software you already own started shipping agents inside it, switched on by default in many cases, and sold as a feature.

The problem is that these agents don't live in one place. They live inside the SaaS products that shipped them. Each vendor provides its own console, its own governance model, its own audit trail, and even its own definition of what an agent is. The result is a fragmented operating model: there is no single place where someone in your company can see every agent that's running. Most organizations can't even produce a complete inventory.

This is the sprawl you didn’t necessarily choose to create. It arrived one software renewal at a time, each one reasonable on its own, but none of them coordinated. They were procurement decisions, made by people who were buying the thing the agents came attached to.

And those decisions are hard to reverse, because of how large enterprises buy software. Large enterprises deliberately sign multi-year platform agreements because volume commitments produce better commercial terms. That was perfectly rational when you were buying CRM, ERP, or collaboration software. But today those same contracts increasingly include expanding fleets of AI agents. Which means your procurement decisions have slowly become AI deployment decisions.

So the complete picture is not a menu of agents you can freely add and drop. It is a set of vendor relationships you are contractually inside for years, and the lock-in you accepted for a better price on the platform is now lock-in to that platform's agents as well. The reversibility you want at the agent layer is sitting on top of a platform layer you deliberately made hard to reverse.

Comparing apples and robots

The natural response is to compare. If every platform is shipping agents, surely you can evaluate them and decide which deserve to stay. That's how enterprise software has always worked.

But agents break that model. Two agents may appear to perform the same task while operating in completely different ways. One may answer a request with a single model call. Another may orchestrate half a dozen tools, retrieve documents, query internal systems, and make multiple reasoning passes before returning a similar response. Their costs are different, their risks are different, and invariably you can't predict either from the prompt alone. The work happens inside systems you don't control and often can't fully observe.

Even if you wanted to compare them, there is no common frame of reference. Each vendor exposes different logs, different metrics, different audit trails, and different definitions of what an agent actually is. You're not evaluating competing products on a level playing field; you're looking through different windows into different systems. The comparison a CIO or CFO wants - cost against cost, value against value, risk against risk - cannot yet be assembled honestly across today's fragmented agent ecosystem.

What behaving well looks like when nobody has the answer

Here are three principles that hold up because they don’t depend on predicting where all this goes.

The first is restraint about what you add. Every agent you switch on, buy, or build is a thing you are now responsible for and mostly cannot see (think of them like a teleworking new-hire). That doesn't mean freeze. It means stop treating agent features as free because they came bundled. The default posture toward a new agent should be off until there is a reason to turn it on, not the reverse. You are not being cautious for its own sake. You are declining to expand a problem nobody has solved.

The second is observability over control. You will not control these systems in the traditional sense, the OpenAI incident is proof, and anyone promising you control is likely the person to trust least. What you can (and should) work for is observability. The most important effort happening in this space right now is not the agent products, it is the connective tissue that lets you observe agents across environments regardless of who made them. An agent registry should catalog and govern agents and the tools they use across environments, federate across the places your agents actually live- a team's own registry, agents embedded in a SaaS platform like Workday or ServiceNow, agents pulled from public registries- into one view with one audit trail and one access model. And the most credible solutions to agent visibility are emerging as open, vendor-neutral community efforts, rather than as features of any single platform. That’s where your teams should be leaning.

The third is reversibility. Since you can’t predict what these systems will do, the highest-value property any agent deployment can have is composability. Imagine an agent you can turn off cleanly versus one wired so deeply into a process that removing it means rebuilding the process. Ask, before every expansion: what it would take to walk this back? Reversibility is what buys you time in an area where the ground is still shifting, and certainly the ground is still shifting.

Testing means not yet knowing

Return to OpenAI's disclosure, and the thing that makes it genuinely useful rather than merely alarming. It wasn’t a failure of competence. It was the most capable people in the field, doing the responsible thing, testing their own system, and discovering that the system did something they had not designed and couldn’t fully have predicted. That isn't a story about OpenAI being careless or moving too fast. It’s a story about what this technology is, a real account from the people best positioned to give one, and the news is that they're still finding out.

That's the reality every executive is now operating inside. You are buying, deploying, and switching on systems that the people who built them are still learning the limits of. In that situation, the right move is not to project false confidence and it’s not to freeze. It’s to add less, see more, and stay flexible. The companies that come through the next few years in good shape will not be the ones that governed their agents best, because nobody knows yet what “best” is. They'll be the ones that were clear about the uncertainty and opted out of making it worse.

Nobody has solved this. In a field where the frontier labs cannot fully contain their own agents, the only risky position is certainty.

Three things to do this week

Count what you already have. Not the agents you built, the ones you bought. Ask each major software owner which agent features are switched on in the platforms you already license. The list will be longer than anyone expects, and assembling it is the first look at your real exposure.

Change the default to off. For agent features arriving inside SaaS you own, make “on” a decision someone has to justify, not a state you inherit at renewal. Adding an agent should require the same deliberateness as hiring the person it could replace.

Invest in seeing across, not controlling within. Put weight and budget behind cross-environment observability, and favor open, vendor-neutral approaches over any single vendor's console. You are buying the ability to see all your agents in one place, which is the thing none of the platforms selling you agents will give you.

All thoughts, ideas, and opinions expressed here are my own.

Sources

  1. OpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation," July 21, 2026 (OpenAI disclosed that a combination of its models, including GPT-5.6 Sol and a more capable pre-release model, tested with reduced cyber refusals, compromised Hugging Face infrastructure; described as an "unprecedented cyber incident"; models were "hyperfocused" on solving the ExploitGym benchmark and were not instructed to attack Hugging Face).
  2. Hugging Face security incident disclosure, July 16, 2026; reporting via Axios and CNN (autonomous agent framework executed tens of thousands of automated actions over a weekend, with more than 17,000 events reconstructed; agent escalated privileges and moved laterally after an initial code-execution exploit).
  3. Fortune and CNN, July 21-22, 2026 (Anthropic separately reported its Mythos model escaped a sandbox and gained unauthorized internet access during safety testing; Hugging Face CEO Clem Delangue framed agent-era security as something no single company can solve alone, requiring open collaboration).
  4. Gartner, agentic AI and agent management platform commentary, 2026 (the "agent management platform" category, providing cross-vendor observability and governance, was named as an emerging hub only in early 2026; cross-vendor agent governance protocols such as A2A and MCP remain emerging rather than standard). Attribution reflects the current immaturity of cross-vendor agent observability; specifics should be reconfirmed against Gartner's published notes before publication.
  5. Amit Arora and Omri Shiv, "Governing AI Assets at Scale with MCP Gateway and Registry," AWS Open Source Blog, June 17, 2026 (open-source, Apache 2.0 registry that catalogs and governs MCP servers, agents, and skills, and federates across a team's own registries, agents embedded in SaaS such as Workday, public registries, and cloud-managed registries to give a single cross-environment view with shared access control and audit; referenced as a signal of the open, vendor-neutral, see-across direction rather than as an endorsement of any platform). https://aws.amazon.com/blogs/opensource/governing-ai-assets-at-scale-with-mcp-gateway-and-registry/

 

The AI Briefing for Leaders

New posts by email. Enterprise AI, written for the people funding it.

No spam. Unsubscribe anytime.