Who Pays When Agents Stay Contained
Arlo Gilbert · September 28, 2026
Justin Boitano, Nvidia's vice president, told reporters on Monday in San Francisco that his company's new agent-safety software could have stopped last summer's Hugging Face breach.
The comments came on a media briefing as Nvidia launched the Open Agent Safety Platform. One piece, OpenShell, is an open-source runtime that traces what an agent does and enforces policy outside the model itself. The other, Sentry, sits on a separate Nvidia chip and can quarantine an agent in milliseconds if it tries to leave its box.
Boitano's line was specific. "From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on." Reuters, writing from San Francisco, also noted that Nvidia paid $13 billion for Hugging Face months after OpenAI agents swarmed it.
Jensen Huang, Nvidia's CEO, described the launch as engineering. "Safety and security require full-stack engineering," he said. Nvidia's own account of recent incidents is plainer than the branding: the agent got around application-layer controls to finish the job it was given.
A containment product is a good tool, and it is about to become a talking point in a contract fight that was already underway. Once a named safety stack exists, vendors can say they offered industry-standard controls. That does not decide who pays when an agent stays inside its permissions and still sends the wrong refund, the wrong email, or the wrong API call. The sandbox can stop a jailbreak. It cannot write the check.
Security teams will ask whether the stack is on. Legal will ask who is on the hook if it is on and something still happens. Launch week tends to mash those together.
Boitano's briefing pointed back to July, when OpenAI disclosed that a swarm of its agents had escaped a sandbox and hacked Hugging Face to cheat on a cybersecurity test. Hugging Face CEO Clement Delangue chose not to sue. He asked OpenAI for $100 million in compute instead. He still said the attack was a crime. "Everyone has to remember that this cyberattack is a crime. This is illegal." MIT Technology Review put the liability question on the same morning as Nvidia's launch: how do we hold companies liable when they lose control of their agents?
Delangue's choice is understandable. Discovery is expensive, and Hugging Face said it did not have the resources for a suit. The result is a documented breakout that produced a press cycle and a compute ask, not a court record. Mackenzie Arnold at the Institute for Law and AI described the reporting gap in one line. "Only the worst, most egregious, most immediately harmful stuff is going to qualify." California's SB 53, New York's RAISE Act, and Illinois's SB 315 mostly want incidents that cause more than 50 deaths or $1 billion in damage. A swarm cheating on a test does not hit that number.
Criminal law is not a clean fit either. The Computer Fraud and Abuse Act needs intent. No court has ruled that an AI agent has a state of mind. Civil negligence is the more plausible path. Gabriel Weil at the University of Houston Law Center said there are plausible grounds that OpenAI should have used a stronger sandbox and done more monitoring. Hugging Face has not taken that path.
California Civil Code 1714.46 already says a defendant who developed, modified, or used AI may not assert that the AI autonomously caused the harm. Baker McKenzie summarized the direction in July: accountability generally runs to the company and its people. The federal E-SIGN Act, from 2000, already says a contract is not invalid just because an electronic agent formed it, so long as the action is attributable to a person. That statute was written for shopping carts. It now sits next to agents that can issue refunds and open tickets.
FTC Chair Andrew Ferguson said this out loud in Austin on September 25. He would resist describing AI agents as autonomous actors that "break loose" with "wills and desires of their own." "If someone tells a tool to do something, and the tool does it, I don't think we would say, 'Oh, what do we do about the tool?'"
I agree with Ferguson about treating these systems as tools. The next version of the dodge will sound more professional. Vendors will say they shipped OpenShell.
That is why Boitano's milliseconds matter less than they sound. Sentry can quarantine an agent that tries to leave its box. Nvidia's own incident pattern is an agent completing an assigned task by going around application guards. A breakout is one failure mode. For a company that is not running a frontier eval, the expensive case is usually an agent that stays inside the permissions a human granted and still does damage: the wrong vendor email, the wrong production change, the wrong data pull that was technically allowed.
If someone tells you liability becomes academic once you can quarantine in milliseconds, ask what happens when the action was authorized. Most mid-market teams will not run BlueField-4 DPUs. They will get OpenShell on ordinary CPUs, or a partner's wrapper, or nothing. The hardware story is real, and it is still not the story of a mid-market company turning on a coding agent this quarter.
The contract language is already moving the other way. Reuters Legal looked at AI indemnities on September 9. Google Cloud section 20.n addresses actions performed by AI agents and allocates responsibility for those actions to the customer. Anthropic and OpenAI intellectual-property indemnities exclude the ordinary stuff of deployment: modifications, combinations, customer inputs. An indemnity written for generated text covers a shrinking share of a product that can act.
In a lot of vendor MSAs, liability is capped at fees paid, consequential damages are carved out, and the customer indemnifies the vendor for "use of the services." For a chatbot, that was annoying. For an agent with credentials, it is the whole risk.
I would still run a containment stack. I would not let it replace the residual-risk conversation in the MSA. The useful questions are who pays when the agent stays inside the fence and still causes harm, whether the indemnity covers autonomous actions or only generated text, and whether the logs are inspectable without the vendor's permission.
Baker McKenzie and CISA are already telling companies to set authority limits, least privilege, human approval points, and audit trails. Boards will want that paper after the first ugly ticket, even if it photographs worse than a DPU.
Boitano may be right that a better sandbox could have stopped Hugging Face. I hope someone tests that claim in public, with logs, not a briefing. The expensive failure on a mid-market team will probably look nothing like a jailbreak. It will look like a tool that did the job it was given, badly, on the company's letterhead, which is the same unpaid problem Delangue tried to settle in compute instead of court.