Skip to main content

What Deloitte's refund tells us about the real agent problem

Arlo Gilbert · August 19, 2026

What Deloitte's refund tells us about the real agent problem

Late in the Australian winter of 2025, Dr. Christopher Rudge, an academic at the University of Sydney, sat down with a report his government had paid Deloitte $440,000 to produce. It reviewed the IT system that the Department of Employment and Workplace Relations uses to automate welfare penalties. Rudge, who follows this area of law professionally, kept noticing that the footnotes did not add up. He found citations pointing to academic papers that had never been written. One of the case references, to the robodebt-adjacent Deanna Amato v Commonwealth, misdescribed what the court had actually decided.

Rudge flagged the errors to the Australian Financial Review that August. By early October, Deloitte had issued a partial refund to the department and quietly reuploaded the report with an updated appendix. The appendix disclosed, for the first time, that portions of the report had been generated with a tool chain built on Azure OpenAI GPT-4o, hosted inside the department's own cloud tenancy. Labor senator Deborah O'Neill put it plainly to the Guardian: "Deloitte has a human intelligence problem." Deloitte kept the recommendations, refunded the last invoice, and moved on.

The $440,000 was almost incidental. The government got its money back, and a dozen faked footnotes are embarrassing rather than existential. What lingers is that a global consulting firm handed a sovereign government a document whose provenance nobody, at any point in the process, was accountable for. No partner was named as the human reviewer. No workflow around the AI outputs was described in the appendix, presumably because there was not enough of one to describe. The AI did what LLMs do when nobody is watching. It wrote confidently and got things wrong.

That is the shape of the problem now facing every enterprise buying agents. The question at the top of the RFP has stopped being "which model." Buyers are leading with a colder set of questions about ownership: who runs the workflow when it breaks, and whether anyone in the building would even notice if it started drifting. The MIT NANDA project's State of AI in Business 2025 study, built on 150 leader interviews and analysis of 300 public deployments, landed on a 95% failure rate for enterprise generative AI pilots. Model quality had almost nothing to do with it. The failures were procurement failures wearing a technical costume.

The arithmetic of the last year, distilled by VentureBeat's Pulse survey of 145 organizations published July 1, 2026, sharpens the picture. Eighty-five percent of respondents now run two or more platforms each claiming to be the "primary" AI layer for their business. Only 38% have a central team that actually governs AI behavior across those platforms. Seventeen percent admit that nobody at their company holds formal accountability for AI at all. Just ten percent have automated monitoring and alerting running against their production AI systems. The rest are relying on human review or on the phone ringing after a customer notices something is off. Deloitte's Australian client noticed the same way.

That was roughly the state of play on June 12, 2026, when a US emergency export-control order pulled Anthropic's Claude Fable 5 offline, worldwide, without warning. Anthropic had no way to verify user nationality in real time, so it suspended access for everyone. Fable 5 had launched three days earlier at $10 per million input tokens and $50 per million output, the smartest and most expensive frontier model on the market. It stayed dark for weeks.

The VentureBeat survey caught the industry mid-flinch. Fifty-one percent of enterprises were already running a hybrid posture, blending closed frontier models for general reasoning with open-weight models on their own hardware for the heavy execution. Another sixteen percent were actively pivoting core workflows onto open weights. The remaining third had gone all in on closed APIs, and they spent June rediscovering what pure vendor dependency feels like. Brian Craig, senior director of architecture at Liberty IT, the engineering arm of Liberty Mutual, described the alternative from the stage of VentureBeat's AI Impact event in New York on June 24, mid-blackout. His company had built what he calls an AI backbone: roughly fifty independently replaceable components for security, governance, observability, and orchestration. "You can't lock in right now in one vendor and even one framework," Craig told the room. "You need to keep being able to have the flexibility."

Craig, who happens to be Irish, watched the export order hit him personally as a foreign national user. He shrugged and kept building, because his procurement decisions three years earlier had assumed something like that morning would eventually arrive.

The buyers I talk to at 500-person companies have started asking a very different set of questions this year. The old RFP asked how a platform would get them access to the frontier models. The new RFP asks whether the vendor will run the first pilot inside the customer's environment for four to eight weeks, and who owns the workflow context and process knowledge that the agent accumulates once we are past the demo. These are ownership questions. Model-quality questions used to lead the meeting; they now sit later in the agenda. They are the questions Deloitte's Australian client should have been asking before the assurance review was commissioned.

At Osano we hear a version of this every week from privacy and security teams. The hard part of shipping an agent is almost never getting the demo to work. The demo is easy. The wall enterprises are running into is legal and security, and eventually ops, all looking at the same deployment and agreeing in writing that they know who is accountable when the agent does what it was designed to do, only wrong. Model access has commoditized. Governance is what still carries a price. Enterprise buyers keep confirming this quietly and increasingly, in their term sheets.

If you are the VP of Ops or the director of engineering signing an AI contract this quarter, the paper needs different language than it did a year ago. Two clauses I would not sign anything without: model portability, so the pilot survives the next Fable 5 morning without a rewrite; and a named human at your own company who owns the deployment end to end, with a public calendar and a job description that contains the word AI. The audit trails follow from that person existing. The most-cited barrier to enterprise AI governance in the VentureBeat data was the absence of that person. In seventeen percent of surveyed enterprises, the role still does not exist.

For the CEO reading this, blaming your engineers for failed pilots is the wrong diagnosis. The pilots keep dying because the org chart cannot answer who holds the leash. Your competitor's chart cannot answer it either, which is why the next round of enterprise procurement will be won by whichever companies fill the governance-shaped hole first. The rest will keep writing Deloitte's appendix.

That appendix is worth reading again for what it leaves out. No partner is named. The internal review that let the fabricated citations survive is not described anywhere, presumably because there was not much of one to describe. The whole document reads as a governance void in the passive voice. The refund covered the invoice, though it did not fix the missing owner. Every enterprise deploying agents right now is one bad Monday away from writing the same appendix. The only real question is whether they will already know the name of the person who has to sign it.