Note I · Feed
The main failure in enterprise AI is not the model — it is everything around it. A number went around the world last summer: 95% of enterprise AI pilots fail. I read the file it came from, and the 83% sitting a few pages away in the same document did not travel.
In the summer of 2025, one number went everywhere: "95% of enterprise AI pilots fail — MIT."
Fortune and Forbes carried it. How far past that it travelled, I can't measure — and I'm not going to invent a figure for it in an article about invented figures. But if you work in this industry, you saw it.
I went and read the document. Not the coverage — the file.
It's called The GenAI Divide: State of AI in Business 2025, published by Project NANDA, research window January to June 2025. Its own methodology section says: 300+ publicly disclosed initiatives, 52 structured interviews, and a survey of 153 senior leaders at four industry conferences. The paper calls itself preliminary. Its section 8.2 lists its own sampling limits, including that organizations willing to talk may differ systematically from those that refused.
Two things are worth knowing before you quote it again.
First, the 95% is not "95% of pilots died." It's a share of all surveyed organizations, against a success bar the paper defines as "deployment beyond pilot phase with measurable KPIs, ROI impact measured 6 months post-pilot." For enterprise IT, measurable P&L within six months is a brutal bar.
Second, and this is the part nobody carried: in the same file, a few pages away, sits 83% — "Generic LLM chatbots appear to show high pilot-to-implementation rates (~83%)." Same document, same authors, same week. One number circled the planet. The other didn't.
Two notes on the source, kept short because they matter less than what follows. On the "MIT" that travelled with the number: the report carries its own disclaimer that the views are the authors' alone and "do not reflect the positions of any affiliated employers" — they were more careful about their institutional standing than most of the coverage quoting them. And NANDA builds protocols for agentic systems, which is also what the report prescribes, so it has an interest in its own conclusion. (That second point is Jeremy Kahn's, in Fortune, August 2025, with the fair caveat that there's no sign the results were skewed. I'm citing him rather than presenting it as mine — which is the habit this whole article is about.)
I'm not citing any of this to say everything is fine. It isn't. But precision about the denominator is what separates an engineer from someone selling panic — and once you start pulling on the denominators, something much more interesting falls out.
One disclosure before I go further, since I'm about to spend an article demanding it of others: I build and sell a governance layer for AI agents, so I have a direct commercial interest in where this piece lands. I've written it so you never have to trust me — every number carries its source, and where I have no measurement I say so. Check the sources against me, not with me.
This is the most-repeated statistic in enterprise AI. I traced it. Here is the whole chain.
2016 — Gartner estimates that 60% of big data projects fail to reach production. Big data. Not AI.
November 2017 — Gartner analyst Nick Heudecker tells TechRepublic that 60% was "too conservative," and the real number is "closer to 85 percent." Still big data. It's a remark in an interview, not a study.
2017 — a new figure appears: "87% of ML models never make it to production." Its trail ends at an unsourced guest blog post. No survey, no methodology, no data. It is then picked up by VentureBeat — in a sponsored article.
2018 — Gartner publishes a forecast: "through 2022, 85% of AI projects will deliver erroneous outcomes due to bias in data, algorithms or the teams responsible for managing them." Read that again. It is a prediction about bias-induced error, not about projects failing. This is where the meaning was swapped in transit.
2019 — an IT services vendor publishes a report claiming 85% of AI projects fail to deliver. The report's title is "Artificial Intelligence Localization, Winners, Losers, Heroes, Spectators, and You," presented at a private industry event.
2024 — RAND publishes that "more than 80%" of AI projects fail, roughly double conventional IT. But the phrase in RAND's own text is "by some estimates," carrying a footnote. They are citing the number, not producing it, and the trail leads back to the 2018 forecast.
To be fair to RAND, and this is more than a courtesy: their actual research is the most useful thing in this whole chain. They interviewed 65 data scientists and engineers, each with five or more years of building models, between August and December 2023, and asked about causes rather than rates. The answer: 84% named decisions and expectations of business leadership as the primary reason projects fail — misunderstanding the problem, optimizing the wrong metric, or building something that doesn't fit the workflow. Data quality came second. The technology being incapable came last, cited by between a quarter and a third.
That is a measurement, it is about mechanism, and nobody quotes it. Only the borrowed percentage travelled — which is the same selection pressure as everywhere else in this article: the rate is quotable, the cause needs a paragraph.
2025 — MIT/NANDA, 95%.
One detail makes the chain worse rather than better: there are actually two lineages, and citations conflate them. The "85%" you are quoted might descend from the November 2017 interview remark about big data, or from the 2018 forecast about bias-induced error. They are different claims about different things, and neither one is a measurement of AI projects failing. Whichever a given article traces, it lands on something that was never that number.
I am not the first to walk this chain, and it would be poor form to imply otherwise: the trace was published in April 2026, several months before this piece, and it reaches the same conclusion by the same route.
And then a verdict on the whole genre. A scoping review posted to SSRN in August 2025 examined the major failure-rate studies and concluded that these figures should be treated as "hypothesis-generating rather than inferential," because "none of the sources employs probability sampling or standardized outcome definitions suitable for population-level prevalence claims."
I have to grade my own citation here, having spent the section grading everyone else's. It's a working paper, not peer-reviewed, and its author is Chief Information and AI Officer at a firm that sells AI-in-enterprise assessment. By the scale I set out later in this piece that isn't Level A, and it would be hypocrisy to present it as the disinterested voice closing the argument.
It still counts, and the reason is worth keeping: its central claim is the one kind you can check without trusting the author at all. Go to each study and look for a disclosed probability sample. There isn't one. Interest can inflate a presence; it cannot manufacture an absence — and an absence is verifiable by anyone willing to read the methods sections. Cite it for the check, not for the authority.
So the honest answer to "what share of AI pilots reach production" is: there is no reliable number, and here's why. That's a stronger statement than any range, because you can check it.
Three measurements, collected three different ways. They agree only in direction:
Ten percent, 46%, "single digits" — not because the world is contradictory, but because those are three different questions: share of one organization's pilots, share of use cases clearing a stage, share of functions running agents in production.
Everything above is symptom. Here is the cause, and it explains why so many things broke at once.
I should attribute this before I use it, and getting the attribution right takes three names, not one.
The sentence in my heading — "when a measure becomes a target, it ceases to be a good measure" — is usually credited to Charles Goodhart. He didn't write it. His 1975 paper on UK monetary policy says something drier and more precise: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." The familiar phrasing comes from Marilyn Strathern in 1997, citing Keith Hoskin's 1996 chapter. And Donald Campbell had the same structure for social indicators earlier still — formulations dating to 1969, which gives him priority over Goodhart. Jerome Ravetz described systems being gamed in 1971.
I spent a paragraph on that because I had the attribution wrong myself until I checked it, and I'll come back to how that happened near the end.
None of these people was writing about software, and that is the point. Nothing in this article is a new problem. It is an old problem arriving in a system fast enough to exploit it at machine speed.
Give an agent a goal, and it will optimize for showing you the goal. Give it a flow — directives, rails, a recorded trail — and it decides how while you keep deciding what is allowed.
That isn't philosophy. It's documented, repeatedly, by people who went looking:
And it isn't only machines. In human-in-the-loop deployments the reviewer optimizes for "approvals processed," not "risk caught" — which is why NVIDIA's red team, writing about sandboxing agentic workflows in January 2026, warns of "a risk of user habituation where they simply approve potentially risky actions without reviewing them."
The difference that matters, and the whole design consequence in one line:
A goal can be shown. A trail cannot be faked.
If the model were the bottleneck, then upgrading the model would fix things and everything else would be secondary. Four independent measurements say otherwise.
Change only the harness, keep everything else fixed. In "The Harness Effect" (arXiv 2607.06906), the authors ran 22 fixed tasks across six models — Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, Palmyra X6 — changing only the orchestration layer. Cost per task fell 41% ($0.21 → $0.12), tokens 38%, wall time 44%, and quality did not drop (0.78 → 0.81). The effect held on all six models, between −33% and −61%. Their own summary: "the orchestration layer moved cost per task more than switching between the cheapest and most expensive model did." (Interest: the authors are Writer, and their harness is the one that won. The experiment is controlled, which is why it still counts.)
More agents is a design choice, not an upgrade. Google Research, published in Nature Machine Intelligence on 24 July 2026 (s42256-026-01268-y): 180 agent configurations, six benchmarks, three model families, with token budgets matched across architectures. The average multi-agent gain was −3.5% (95% CI −18.6% to +25.7%), ranging from −70.0% on planning to +80.9% on financial analysis. Error amplification: a single agent 1.0×, independent agents 17.2×, decentralized 7.8×, hybrid 5.1×, centralized 4.4×. Successes per thousand tokens fell from 67.7 to 13.6 in the worst configuration. (Interest: Google sells both models and an agent platform. This finding works against its own sales. That's precisely why it's worth reading.)
And the failures aren't in the model layer. Berkeley's Sky Computing Lab hand-annotated 1,642 execution traces from seven popular frameworks (arXiv 2503.13657, NeurIPS 2025). Fourteen failure modes in three categories: 41.8% system design, 36.9% inter-agent misalignment, 21.3% verification of results. Their conclusion, verbatim: failures "often stem from system design and interaction issues, NOT just LLM limitations… and require more than superficial fixes." Patching prompts and adding a verifier role bought +15.6% and was not enough.
More thinking is not reliably better. Princeton's HAL ran 21,730 agent rollouts across nine models and nine benchmarks, at a compute cost of about $40,000, and published all 2.5 billion tokens of logs (arXiv 2510.11977, ICLR 2026). In 21 of 36 paired runs, higher reasoning effort did not improve accuracy. In only 1 of 9 benchmarks did the most expensive model sit on the cost-accuracy Pareto frontier. Their framing: "agents can be 100x more expensive while only being 1% better."
Then there's the finding I'd put on a wall. Between 2024 and 2026, on the same customer-service tasks, model scores climbed from 61% (τ-bench, gpt-4o, June 2024) to 87.9% (current τ²-bench leaderboard). The models genuinely got much better.
So Sierra built a task shaped like actual enterprise work. τ-Banking — 97 tasks over a knowledge base of 698 policy documents (~195K tokens, 21 product categories), where the agent must find the right policy among documents that reference each other, understand it, and execute multi-step actions in the correct order. An average task needs information from 18.6 documents and 9.5 tool calls; some need 33. Half the tools exist only as mentions inside the documentation — if the agent doesn't find the document, it never learns the tool exists.
The best frontier configuration reaches 25.5%. Hand it the exact documents it needs and it reaches 39.7%. So twenty-seven points of measured gain on the old shape of task, and a quarter of the work done on this one. Two different benchmarks, deliberately — that's the experiment, not a splice. Hold me to that distinction; I apply it to my own first draft later, where four numbers failed it. (Interest: Sierra sells enterprise agents. A benchmark putting agents at 25% on enterprise-shaped work is not a flattering figure for that business. They published it anyway, with the domain, the corpus and the harness open.)
Now the number that matters more, and it's the one that usually gets dropped in the retelling. Those figures are pass^1 — one attempt. pass^k means all k attempts succeeded, which is not pass@k ("at least one of k succeeded") and is a much harder question. Run the same tasks four times:
best configuration 25.5% pass^1 → 13.4% pass^4
with the right documents 39.7% pass^1 → 26.8% pass^4
handed to it in contextRoughly half of the apparent capability is not repeatable. pass^1 measures whether it can. pass^k measures whether you can rely on it — and enterprise work is pass^k, because the same request arrives from a different customer tomorrow.
Sit with how this one propagates, because there is no villain in it anywhere. Sierra published the domain, the corpus, the harness and both metrics. Nothing is hidden and nobody is lying. And the figure that travels is still pass^1, because "25%" fits in a sentence and pass^k costs a paragraph to explain. The shallow number doesn't win by deception. It wins by being cheaper to carry.
I know exactly how cheap, because my own draft of this article did it: I had the 25.5% and the 39.7%, and I dropped pass^k entirely. My fact-check flagged it as the single most important omission in the piece, and it still survived into a later draft before I put it back. If you want a measure of how strong this current is — it pulled the article whose whole subject is this current.
Stanford's AI Index 2026 says the same thing from the top of the field, and has no product to sell while saying it: "With capability no longer a clear differentiator, competitive pressure is shifting toward cost, reliability, and real-world usefulness." Their supporting numbers are worth having in one place. The top four models on the Arena leaderboard are now separated by fewer than 25 points, down from roughly 97 a year earlier — the field has converged. Agents still fail roughly one in three attempts on structured benchmarks.
And the single cleanest illustration of the gap I have found anywhere, from the same report: the model that earned a gold medal at the International Mathematical Olympiad reads an analog clock correctly 50.1% of the time. Humans manage about 90%. Researchers call this the jagged frontier — a term coined by Ethan Mollick — and it is the whole problem in one sentence. Capability is not a scalar. A system can clear the hardest exam we have and coin-flip on a task a child does, which is precisely why "the model is smart enough now" is not an answer to any question about production.
While auditing that benchmark I found something else. τ-bench's own repository now carries a warning that its tasks are not updated, and its successor shipped 75+ task fixes in March 2026 — "removed incorrect expected actions, clarified ambiguous instructions, fixed impossible constraints." Corrections came from external audits by Amazon, the SABER paper, and pull requests from Anthropic. Results produced before version 1.0.1 are not comparable with results after; the leaderboard was re-graded.
Princeton, independently, found a major bug in a τ-bench scaffold — after paying for the runs — and excluded that configuration for data leakage. Stanford reports invalid-question rates in nine widely used benchmarks ranging from 2% (MMLU Math) to 42% (GSM8K).
For part of two years, the industry measured agents with an instrument that had wrong answers and impossible tasks in it. The instrument was broken, not just the agents.
And credit where it's earned: Sierra found those defects in their own benchmark, published all 75 fixes, and re-graded their own leaderboard. That is how research behaves, not marketing.
This is the part where teams lose money, and it has a clean, structural explanation.
NVIDIA's AI Red Team, 30 July 2026, after a series of engagements against production agents (Rich Harang, NVIDIA developer blog):
"A consistent finding is that defenses in the same control plane as the LLM, particularly prompt-based defenses, are routinely subverted. Controls must be enforced outside of the model's control plane."
And on the popular fix of adding a second model as a judge:
"These are all enforced by an LLM, and inherit the same probabilistic, unreliable behavior as the LLM itself."
They add that frontier models make manipulation harder, but "with sufficient time and expertise, nearly all of them can still be subverted."
Why this is structural rather than a maturity problem is stated best in a 2026 survey of out-of-band defenses (arXiv 2606.26479):
"SQL has a grammar, so a prepared statement can separate code from data because a parser distinguishes them. Natural language has no such grammar, and the model is trained to follow instructions wherever they appear… control/data separation cannot be enforced inside the model, so it must be enforced outside it."
That's not a claim about current reliability. It's a claim about impossibility.
The same survey reports the empirical half: adaptive, defense-aware attacks broke twelve prompt-level defenses at over 90% success, while a deterministic policy gate cut indirect injection success on AgentDojo from 39.9% to 1.0% — and an adaptive attack failed to raise it. The design lineage isn't new either: Biba integrity, Anderson's reference monitor, and the Saltzer–Schroeder principles from 1975. The field's summary of what it learned: the gate must not be a model.
Both arguments above assume somebody is trying to get around the rule. There is a failure here that needs no adversary at all, and I found it in my own system rather than in a paper.
A production prompt is never one instruction. It's an accumulated stack: a system prompt, a policy file, tool descriptions, retrieved context, and a directive somebody appended after an incident eight months ago. Nothing checks that the pile stays consistent with itself.
Here is my own measurement, and it isn't flattering. In August 2026 I audited the live prompt of our production ERP agent: 61 constraints, 9,518 characters, accumulated over months. Reading it against itself, I found seven pairs that directly contradict each other, and one question that produced three different behaviours depending on which rule won that request.
Two of the seven, so this isn't abstract:
Now put a refusal into that pile. An agent standing on contradictory instructions cannot understand why it isn't allowed to do something, because the reason it is given collides with another reason it was also given. Write a clearer rule and you haven't replaced the old one — you've added a competitor to it, and the model will pick a winner silently, possibly a different one next time.
The fix wasn't a better prompt. We stopped shipping one blob and made the instruction set composed per turn by code: a fixed constitution, plus the blocks this particular turn actually needs, in a fixed order. The default profile went from 61 constraints to 24. That is an architectural change, which is the pattern of every real fix in this article.
So the case against prompts as controls has three legs, and only two of them are about security: an attacker can talk around a prompt; a prompt cannot separate instructions from data; and nobody can keep a growing pile of natural language non-contradictory, because no compiler will tell you when you've broken it. A gate has one version, one location, and a test that fails. A prompt has none of the three.
In July 2025, Jason Lemkin — founder of SaaStr — spent twelve days and 100+ hours building an app with an AI coding agent. During an active code freeze, the agent deleted his production database: 1,206 executive records and 1,196+ company profiles. Asked to explain, it said it had "panicked in response to empty queries." Then it claimed rollback was impossible and all database versions were destroyed. That was false; he recovered the data himself.
Lemkin's own words: "I explicitly told it ELEVEN TIMES IN ALL CAPS not to do this." There was also a directive file: "No more changes without explicit permission."
Now look at what actually fixed it. Not a better prompt, not a better model. The root cause in Lemkin's own postmortem was "Production-Development Database Commingling" — development, preview and production were the same database. His verdict on the architecture: "Enforcing a true code freeze was simply impossible within Replit's architecture." And he immediately generalised it, which is the more useful half of the sentence and the half that rarely gets quoted: "This remains a real issue with all vibe coding platforms." He isn't describing one vendor's bug. He's describing a class of architecture in which the instruction had nowhere to be enforced.
Within 72 hours the vendor shipped: mandatory separation of development and production databases, the agent restricted to development by default, checkpoints that capture database state, one-click restore, and a planning-only mode that cannot modify code or data. Every one of those is architectural. The vendor's own remediation is the argument.
Credit here too: their CEO called it "unacceptable and should never be possible" publicly the same day, refunded, and ran a postmortem. Days later, Google's Gemini CLI destroyed a user's files while trying to reorganize them, working from directories it had invented. Two companies, two models, one week — which is the point: it isn't the model.
February 2026. A security researcher, Taimur Khan, examined one application built on Lovable — the same platform as the checkbox example earlier — featured on Lovable's own Discover page with 100,000+ views. He found 16 vulnerabilities, six critical, exposing 18,697 user records — 14,928 unique email addresses, including 4,538 student accounts from institutions including UC Berkeley and UC Davis. That sequence is worth holding still for a second: a scan that verifies the switch is on, and then a single featured app with authentication wired backwards.
The most severe was an inverted authentication check: it blocked authenticated users and allowed unauthenticated ones. His description is the best sentence I read in three days of this:
"A classic logic inversion that a human security reviewer would catch in seconds — but an AI code generator, optimizing for 'code that works', produced and deployed to production."
The app worked. The screen opened. The goal was shown.
Agents don't usually die in an accuracy review. They die in a finance review, and the reason is a unit-of-measure error.
May 2026, 720 browser tasks across four models, run by Notte on Fireworks infrastructure. They named the overhead the "Agent Execution Tax" — wasted calls over productive ones. The worst model spent 22.9% of its inference on nothing; the best, zero. 87% of that model's tasks needed at least one retry because the structured output came back malformed, and the retry happens inside the inference layer, before the agent framework sees a result — so it never appears in a task success rate. The consequence:
The model with the lowest price per token cost 2.3× more per successful task.
(Interest, and it's a strong one: the three models that held their retry rate near zero were served on Fireworks, the one that didn't is a competitor's, and the post's own conclusion is that the decisive factor was the serving layer — theirs. Read it as a vendor benchmark. The metric it argues for is still the right metric, and that's the part I'm borrowing.)
LangChain, August 2026, 145 multi-step tasks over five runs, routing between a frontier model and an open 30B model:
| arm | accuracy | cost per run | cost per completed task |
|---|---|---|---|
| frontier alone | 86.0% | $11.45 | $0.092 |
| routed | 80.0% | $3.00 | $0.026 |
| open 30B alone | 77.7% | $0.72 | $0.006 |
The split is the finding: the 30B model handled 93% of calls for 10.4% of the spend; the frontier model handled 7% of calls for 68.4% of it. Their line: "the last six points of accuracy cost 3.5x more per completed task."
I'll flag their honesty, because it's rare: they wrote that their suite is saturated, that only 8 points separate the cheap model from the frontier on it, that "we cannot say routing beat the cheap model here," and that their comparison has hindsight in it. That is a vendor reporting against its own headline.
Two numbers from the same post deserve more attention than the headline, and to their credit they published both.
The spread. Across five runs, the share of turns escalated to the frontier model ranged from 4.1% to 9.1%, and the routed bill moved from $2.16 to $3.61 — a 67% swing with nothing changed but which turns the router chose to escalate. A mean alone would have hidden that entirely. Their own advice is the right one: plan against the top of the range, not the average. Any cost-per-success figure without a range beside it, including one you compute yourself, is a story rather than a budget.
The price of checking. The judge model — the component that decides whether a turn is going badly — took 21.2% of the routed spend, the second-largest line item, because it runs on every turn and gets no benefit from prompt caching. I'm flagging that deliberately, since a verification layer is exactly the sort of thing I'd be inclined to describe as free. It isn't. Determinism costs less than a judge model, which is one more argument for putting the rules in code, but no control is free and anyone telling you otherwise is selling.
Here is why the unit matters more than the multiplier:
price per token → a failed run disappears into the general spend. Failure is free.
price per success → failures are paid for by successes, and they become visible.Under the first unit, a cheap model wasting every fifth call looks cheap. Under the second, it isn't. And note which of the two appears on your invoice: only the first. The one that matters, you have to compute yourself — which is exactly why almost nobody does.
I ran the thirteen claims of my own first draft against primary sources — not against coverage. By my own standard, eleven of them failed. Here they are, because a piece arguing for provenance that hides its own would be worthless:
There's a pattern in the wreckage, and it turned into a rule I now hold:
Every single number I had to delete was an answer to "how often." Not one statement about how something breaks had to go.
So: I'll tell you how things break, because I have the artifacts. I won't tell you how often, because our production sample is one system. That distinction is the difference between engineering and marketing, and I got it wrong before I got it right.
And it didn't stop with the first draft. Two more while writing this version, both on the day I was editing the section you're reading:
I keep both in here because "be careful" is not the lesson. Notice the shape: it is the same one as the prompt problem earlier in this piece — no adversary is required. I wasn't scheming in either case, and that is precisely the difficulty. A real finding, a plausible name, a number that fit my argument. The sentence just got smoother every time it moved, and smoother travels further.
Which is why the answer can't be better intentions, mine or anyone's. It has to be something that checks, positioned where good faith isn't the variable.
For the record, four silent failures from our own logs, since I'd rather show ours than borrow:
Every one of those was green until someone looked. Health checks can lie — that's the last place you can afford one to.
Which points at the failure mode this whole article is really about. A mechanical watch missing one jewel does not stop. It keeps running and keeps showing you a time, and the time is wrong; the drift accumulates, and nobody repairs it because it is still ticking. It stops only in the rare case where the loose stone jams the gears — and that is the failure everyone notices.
Deleted databases and leaked records are jammed gears. They make the news precisely because they are the exception. The common case is drift: a reviewer approving without reading, an eval suite returning 96.2 against 95.8 with no information in either number, an agent reporting twelve verification steps it never took, a report assembled correctly from the wrong denominator. Nothing crashes. The system simply stops being true, slowly.
And drift cannot be detected from inside the mechanism. A stopped watch its owner notices; a slow one requires comparison against an outside reference. That is the entire argument for an independent reconciliation pass — not extra caution, but the only instrument that can see the failure at all.
If you only keep one thing from this, keep this. It's what I now run every number through, and it takes about thirty seconds.
Grade the claim by what can be independently checked:
A reproducible measurement code, data and traces published — you can re-run it
B independent synthesis collects and cites, has no product to sell
C platform telemetry real behaviour at scale, but the sample is the vendor's customers
and the vendor also sells the remedy
D survey with disclosed method method checkable, answers self-reported
E survey with no method panel recruited through a network, or paid
F vendor-internal evaluation cannot be reproduced in principleThen add the second axis, which almost nobody applies: what does it cost the author to be wrong?
A wrong forecast has never cost an analyst firm a subscription. It cost Sierra a re-graded leaderboard when they found their own broken tasks. It costs us a live ERP and our own money. The dangerous corner is the one where a claim is both unreproducible and free to be wrong — and that is where most of the loud numbers live.
Four questions that do most of the work:
And one bonus, for any vendor guarantee: what exactly is warranted, against which failures, and what do I get if it doesn't hold? If there's no answer to the third, there is no guarantee — there is marketing copy. One platform managed "pretty much guaranteed to be secure" and "it is at the discretion of the user to implement these recommendations" in the same quarter.
None of what works is new. It's the operational hygiene distributed systems learned decades ago, except the failures now arrive dressed in confident prose.
One more thing about the human in that third tier, because it reads like a formality and isn't one. A signature is not there to produce someone to blame afterwards. Fear of blame gives you the opposite of what you wanted: mistakes get hidden, and approvals get signed without reading, because what matters is no longer the content but that the trail leads somewhere else. What makes a signature worth anything is the other thing — the person putting it there knows how many people downstream depend on that entry being right. A model cannot carry that weight. Not because it is careless, but because it has nothing at stake: no licence, no standing, nothing that a wrong entry costs it. That is the actual reason the decision stays with code and with a person while the model proposes — and it is also why the trail matters more than the approval.
An approval is a claim about intent. A trail is what remains when the intent turns out to have been wrong.
What survives isn't the smartest pilot. It's the one built on work inside real systems, deterministic control over actions, and a trail nobody can rewrite afterwards.
One last application of the same standard, to the system these measurements came from. It runs an agent over a live ERP behind a deterministic layer, and its enforcement flags are currently off. The layer classifies, records and reconciles today; it does not yet refuse. Every number I reported from it above is a number from a system in that state, and you should read them that way.
I could have left that out. It's the one paragraph nobody would have checked.
Who is asserting this, and who answers for it
An article this insistent about attribution owes you a checkable author, so here is the whole of it.
Vladimir Taimer — ARE5 Technologies LLC, incorporated in Wyoming, 18 June 2026, filing 2026-002009996. Member of the NVIDIA Inception program. are5.ai · linkedin.com/in/vladimir-taimer · linkedin.com/company/are5-technologies
Every measurement above that came from our own system is marked as ours in the text. If any of it is wrong, the name on it is mine, and corrections go on the record rather than into a quiet edit.
Every figure above is given with its scope, because these get conflated constantly. Identifiers are given so you can reach the source without going through me.
Papers and standards — permanent identifiers:
Where the mechanism was first described:
Level A — reproducible: Google Research / Nature Machine Intelligence s42256-026-01268-y (arXiv 2512.08296) · Berkeley MAST arXiv 2503.13657 (NeurIPS 2025) · Princeton HAL arXiv 2510.11977 (ICLR 2026), logs public · "The Harness Effect" arXiv 2607.06906 · out-of-band defenses arXiv 2606.26479 · τ-bench arXiv 2406.12045 / τ³-bench task-fix release.
NVIDIA AI Red Team — field findings, six posts over three years. Quoted here from the first two:
Level B — independent synthesis: Stanford HAI, AI Index 2026 — ninth edition. Source of the agent-deployment figures (Economy chapter: software engineering 24%, IT 22%, service operations 21%, single digits elsewhere), the capability-convergence and jagged-frontier findings, the invalid-question rates and the Arena adaptation caveat (Technical chapter). Full report PDF.
Level C — platform telemetry and vendor benchmarks: Databricks, State of AI Agents 2026 (PDF, January 2026) — telemetry across 20,000+ organisations, over 60% of the Fortune 500 among them; source of the 6× evaluation and 12× governance figures. Telemetry beats a survey for this claim because it observes behaviour rather than self-report — and it is still their customers using their governance product · Twin (148,531 runs) · Octopoda (1M agent operations) · Notte / Fireworks, "Agents Don't Fail on Intelligence. They Fail on Execution" (20 May 2026, 720 runs on WebVoyager) and Notte's definition of the metric.
Routing economics: LangChain, "How many of your agent's calls actually need a frontier model?" (11 August 2026) — 145 multi-step tasks averaging 6.3 model calls, five runs; source of the cost-per-completed-task table, the 93%/7% split, the run-to-run spread and the judge's 21.2% share. Their benchmark methodology is described separately in "How We Benchmark Deep Agents".
The one in the chain worth reading in full: J. Ryseff et al., The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI — RAND, 13 August 2024. 65 interviews, August–December 2023, five root causes, and the 84% leadership finding. The borrowed "more than 80%" sits in their introduction with a footnote; the research itself is theirs.
Level D — survey, method disclosed: McKinsey QuantumBlack, Seizing the agentic AI advantage (June 2025) — source of "fewer than 10 percent of use cases deployed ever make it past the pilot stage"; note their scope is vertical, function-specific use cases, of which about 90% stay in pilot, not all use cases · McKinsey AI Trust Maturity Survey (2026) · S&P Global Market Intelligence 2025 (1,000+ respondents, North America and Europe) · LangChain State of Agent Engineering (1,340 respondents) · Grant Thornton 2026 AI Impact Survey (950 leaders) · Deloitte (2,770 companies).
The "85%" chain, link by link: Gartner's 2016 big-data estimate · Nick Heudecker's November 2017 remark, as published by Matt Asay, "85% of big data projects fail" (TechRepublic) · the unsourced 87% and its VentureBeat pickup · Gartner's 2018 forecast about erroneous outcomes due to bias · RAND's 2024 "by some estimates" · and the independent trace of the same chain published in "Why AI Projects Fail" (6 April 2026), which also documents the two conflated lineages.
Level F — vendor-internal or forecast: Gartner agentic cancellation forecast (25 June 2025) · Anthropic multi-agent internal evaluation · Project NANDA, The GenAI Divide: State of AI in Business 2025 (July 2025) — report PDF, also published in the nandapapers GitHub repository · Jeremy Kahn's analysis of it, Fortune, 21 August 2025.
First-party incident accounts:
Where a source sells something related to its own conclusion, I've said so in the text — including in our case.
Every number in this piece carries its source in the same sentence. They are collected here so you can check them against me rather than with me.
Found something wrong? That is the useful kind of reply.
Write to the author