Some of the most consequential AI failures have very little to do with the model itself.
A classifier tests well and gets quietly disabled after launch because nobody defined who was accountable when it was wrong. Gartner has tracked a version of this at the industry level. In 2024, they predicted that by the end of 2025, at least 30% of generative AI projects would be abandoned after the proof of concept. That wasn’t about weak models. It was about poor data quality, weak risk controls, creeping costs, and no agreement on what business value was even supposed to look like.
A retrieval system answers the question correctly but adds a review step nobody has time for. That same failure mode shows up in MIT’s NANDA research, at enterprise scale. Studying more than 300 enterprise AI initiatives through mid-2025, they found a “shadow AI economy”: employees quietly routing around the sanctioned tool because the official workflow carried more friction than it was worth. Across those same initiatives, 95% of generative AI pilots delivered zero measurable P&L impact, on top of $30 to $40 billion in enterprise investment.
An agent automates the task perfectly and somehow makes the job worse because the person doing the work now spends the day checking the agent instead. BetterUp Labs and Stanford’s Social Media Lab put a number on exactly this in 2025. They surveyed 1,150 full-time desk workers and found 40% had received AI-generated “workslop” — output that looks finished but isn’t — in the past month alone. Each incident took about two hours to clean up. Harvard Business Review covered the same research and did the math: roughly $186 per employee per month, close to $9 million a year at a 10,000-person company. That’s not a model problem. That’s the cost of a system that hands off unfinished work and calls it done.
Those are not problems a bigger model fixes. And better prompting does not solve them. We spend a lot of time asking which model is smartest, cheapest, fastest, or hallucinates least. Those are real questions.
MIT Sloan Management Review and BCG’s 2024 “Winning With AI” research found the same gap, at scale. Roughly 7 in 10 companies report minimal or no measurable business impact from their AI investment. Of the roughly 90% of organizations that have invested in AI at all, fewer than 2 in 5 report any business gain. They just aren’t always the first questions that determine whether an AI system will create value.
AI value isn’t a property of the model. It’s an outcome of the system around it.
I think about that system across five dimensions.
The model gets the attention. Hover a node for detail, or click to jump to what surrounds it.
People
Who owns the outcome? Where does human judgment still matter? Who has the authority to intervene when the system is wrong? I ask these three questions on every project. Most of the time, nobody in the room has a real answer for any of them.
That’s not just my own experience. McKinsey ran the numbers in 2025 — a global survey of nearly 1,500 executives across 101 countries. Fifty-one percent had already had something go wrong with their AI. A bad call. A wrong output. Real damage. And still, only 28% said their CEO takes direct responsibility for governing it. Only 17% said the board does. Ownership thins out before it even reaches the top.
Lower down, it doesn’t disappear. It just changes shape. MIT Sloan’s research on responsible AI found the same pattern inside individual teams: accountability for AI is widely shared and rarely owned. One governance lead they interviewed said it plainly — if it’s everyone’s job, it’s no one’s job. Their researchers even found fairness reviews run by the same team that built the model. Nobody outside the room was checking the room. I don’t call that oversight.
So who actually has the authority to step in when the system gets it wrong? On paper, someone. In practice, often no one. Deloitte found that only 21% of enterprises have a mature governance model for their AI agents. That means roughly four out of five don’t have a clear line between what an agent can decide on its own and what needs a human to sign off — no real-time monitoring, no audit trail. And when nobody draws that line on purpose, something else draws it for you. Leaders who don’t explicitly assign decision rights don’t get a system that waits for instructions — they get one that assumes them anyway, starts setting its own priorities, and starts making trade-offs nobody agreed to. I’ve seen it happen — a default nobody chose becomes the policy nobody questions.
Here’s what all of that adds up to. Gartner polled more than 3,400 organizations already running agentic AI, and their forecast isn’t a crash — it’s an abandonment. More than 40% of agentic AI projects will be canceled by the end of 2027. About 40% of enterprises will quietly demote or shut down their AI agents, and the reason is almost always the same: the governance gap wasn’t found until after something had already broken in production. That’s not a fix. That’s a quiet decommission. MIT’s Project NANDA saw the same shape from a different angle. They reviewed more than 300 enterprise AI programs, and despite $30 to $40 billion in investment, 95% of the pilots delivered no measurable financial return. The models weren’t the problem. Nobody owned the workflow. Nobody owned the feedback loop. The tool never got adapted to the job, so the pilot just got shelved.
That’s the pattern I keep running into. A system without clear ownership does not necessarily fail loudly. It doesn’t crash. It doesn’t set off an alarm. It just quietly stops being trusted — one skipped review, one silent workaround, one shelved pilot at a time.
Work
Is AI improving the work, or simply automating the workflow that already exists?
Those are different things.
Automating a poor process can make the poor process happen faster.
Michael Hammer named this problem in Harvard Business Review back in 1990. He called the piece “Don’t Automate, Obliterate.” His point: lay new technology on top of a broken process and you just execute the flaw faster. His fix wasn’t automation. It was redesign. Ford didn’t automate their old accounts-payable process. They rebuilt it, and headcount dropped from 400 to 5.
Most companies still aren’t doing that. MIT researchers studied over 300 public AI deployments in 2025 and talked to more than 50 leaders running them. Ninety-five percent of the pilots produced no measurable P&L impact. Not because the models were bad. Because the tools got bolted onto workflows nobody bothered to redesign.
McKinsey found the other side of that same coin. In their 2025 global AI survey, the companies actually seeing profit impact from AI were 2.8 times more likely to have redesigned their workflows around it than the companies that didn’t — 55 percent versus 20 percent. Out of 25 organizational factors McKinsey tested, workflow redesign moved the needle more than anything else.
The more interesting question is sometimes:
Should this process exist at all?
Architecture
What actually needs to surround the model?
Retrieval. Memory. Tools. Orchestration. Evaluation. Deterministic logic. Sometimes another model. Sometimes less AI.
None of that is a hunch. Berkeley’s AI research lab studied it in 2024. They gave it a name: compound systems. Their finding was direct. State-of-the-art results increasingly come from systems built out of multiple components. Not from a single model call. They showed the receipts, too. AlphaCode 2 doesn’t trust one answer. It generates up to a million candidate solutions. Then it filters them by actually running the code. That’s deterministic logic doing work a model can’t do alone. Medprompt beats specialized medical models the same way. Not by being bigger. By chaining GPT-4 to retrieval, chain-of-thought, and ensemble voting. AlphaGeometry pairs an LLM with a symbolic solver. Sometimes another model. Sometimes math instead of a model.
Retrieval and tools aren’t a fringe pattern. They’re the default. Databricks pulled usage data from more than ten thousand customers — over 300 of them Fortune 500 — between February 2023 and March 2024. Seventy percent of the companies using generative AI had already wired in tools, retrieval systems, and vector databases. They weren’t waiting on a bigger model. They were building a system around the one they had.
Orchestration is where the gap shows up hardest. McKinsey ran a survey in 2025. Sixty-two percent of organizations said they were experimenting with AI agents. Fewer than ten percent had scaled one into an actual business function. That’s the whole essay in two numbers. Everyone can make the model call. Almost nobody has built the orchestration that turns it into something you can trust at scale.
Gartner named the missing layer outright: agentic AI infrastructure “embeds deterministic guardrails, orchestrates autonomous workflows and provides deep observability at machine speed.” Guardrails. Orchestration. Observability. Three separate jobs. None of them is the model.
Memory and evaluation belong on this list too. I’ll be straight about them. I don’t have a study that isolates memory the way Berkeley isolated retrieval or Databricks isolated tools. I don’t have one for evaluation either. Nothing that measures it as its own category, separate from the guardrails and observability Gartner just described. Nobody’s published those numbers yet. I still build both into every system I ship. That’s an assertion from what I’ve built. Not a citation.
Same with “sometimes less AI.” No one’s published a number showing that swapping a model call for a deterministic rule measurably improves a system. I believe it anyway. I’ve done it. I’ve watched the system get more reliable, not less capable. That’s experience talking. Not evidence.
The model call may be the most visible component. It is rarely the entire system.
Governance
Who approves the use case?
Who owns the risk?
What evidence is required before it ships?
How is performance monitored?
When does a human need to remain accountable?
And who has the authority to pause or retire the system?
I used to think the first two were rhetorical. They’re not — NIST wrote the answer into its AI Risk Management Framework: split the team that builds the system from the team that tests it, and write down who owns AI risk at every level, not just at the top. Senior leadership is required to declare the company’s actual risk tolerance. That’s NIST’s bar. Mine is stricter: a real decision, not a slide deck nobody reads twice.
Most companies skip the structure part entirely, and the scale of that surprised me. McKinsey surveyed nearly 1,500 executives across 101 countries in 2025 and found only 28% have their CEO overseeing AI governance. Just 17% have their board doing it. On average, a company names exactly two people to govern AI — full stop — while the risks those two people are supposed to track have doubled since 2022, from an average of two to an average of four. Same two people. Twice the risk.
Ask the harder question — who can actually pull the plug — and it gets worse. MIT Sloan Management Review put it to AI leaders directly: when a model does something it shouldn’t, who has the authority to stop it? Most couldn’t answer. Their diagnosis matches what I’ve watched happen in practice: chief AI ethics officers and titles like it are, in the piece’s own words, “completely necessary but merely advisory.” Advisory doesn’t decide anything. Real accountability is one named person who can say no to a deployment, or shut it down, and have that stick.
Monitoring is the part people treat as a checkbox. It can’t be. Governance has to run continuously, at runtime, for the whole life of the system — not once before launch and never touched again. Informal, one-off controls stop working the moment AI moves into production at scale.
I don’t think that’s a coincidence. Stanford’s 2026 AI Index counted 362 documented AI incidents in 2025, up from 233 the year before. Governance is catching up, a little — dedicated AI governance roles grew 17% that year, and the share of businesses with no responsible-AI policy at all dropped from 24% to 11%. But the same report names why the rest are still behind: knowledge gaps, budget, regulatory uncertainty. Governance is maturing. It’s just maturing slower than deployment is.
A guardrail can block an unsafe input.
It cannot answer those questions.
Economics
Does the value survive the real cost of operating the system?
Not just inference cost.
Human review. Maintenance. Monitoring. Integration. Retraining. Exceptions. Governance. Change management.
The first time I saw the real number, I assumed it was a fluke. MIT’s Project NANDA looked at 300-plus disclosed enterprise AI deployments in 2025 — 52 structured interviews, 153 survey responses — and found that enterprises had already sunk $30 to $40 billion into generative AI pilots. 95% of those pilots showed no measurable P&L impact. Only 5% made it from pilot to real financial value. It wasn’t a fluke. The report’s own conclusion wasn’t model quality. It was integration and workflow adaptation — the operating cost of actually running the thing inside a business.
Then McKinsey’s 2025 State of AI survey landed, and it was the same shape from a completely different data set. 88% of organizations use AI regularly in at least one function. Only 39% can point to any enterprise-level EBIT impact. Just 6% call themselves AI high performers — 5% or more of EBIT tied to AI. McKinsey tested 25 organizational factors trying to explain the gap. The biggest one wasn’t which model anyone picked. It was whether they’d redesigned the workflow around it.
BCG went further and actually priced out where the value dies. They surveyed over 1,000 senior executives across 59 countries in 2024 and found 74% had nothing tangible to show for their AI investment. When they broke down the cause, about 70% of it was people and process — change management, workflow redesign, governance. Only 20% was technology and data. Just 10% was the algorithm itself. Three different firms, three different survey populations, and the ratio kept landing in roughly the same place: the model is rarely the problem. The system around it is.
Gartner had already called this coming in 2024, before any of those numbers existed: at least 30% of generative AI projects would be abandoned after proof-of-concept by the end of 2025. Not because the models failed. Poor data quality. Weak risk controls. Costs that kept climbing. No clear business case. Their VP analyst Rita Sallam put it in one line I still think about: “As the scope of initiatives widen, the financial burden of developing and deploying GenAI models is increasingly felt.”
Four research teams, four separate populations, and I keep landing on the same read: the model was rarely why these projects died. An AI system can work technically and still fail economically.
These dimensions are not separate boxes.
The interesting decisions usually happen between them.
Compute tiering is both an architecture decision and an economics decision.
A shared ontology is an architecture choice that changes how people, agents, and systems interpret the same information.
Falsifiable evaluation affects both technical reliability and governance.
A Gauntlet Loop that generates, challenges, revises, and re‑tests an answer changes the architecture, but it also changes what we are willing to trust enough to ship.
That is why I use this framework when I build.
Not because five dimensions look good on a slide.
Because they change the questions I ask.
Instead of:
Which model should we use?
I start asking:
Who owns this when it’s wrong?
Are we improving the work or accelerating the old workflow?
Does the architecture fit the problem?
What evidence do we need before this goes live?
Does the value survive the cost of running it?
Those questions tell me considerably more about the eventual system than a benchmark alone ever could.
That is also how this site is organized.
You will find agents, evaluation, memory, knowledge graphs, context engineering, security, governance, and model routing here.
But those are not the point.
They are pieces of a larger system.
The model is one part of it.
Durable value depends on what surrounds it.
That’s what I write about here.
And it’s what I build to test.
Sources
- MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025,” as reported by Fortune. fortune.com
- Gartner, “Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025.” gartner.com
- BetterUp Labs & Stanford Social Media Lab, “Workslop.” betterup.com
- Harvard Business Review, “AI-Generated ‘Workslop’ Is Destroying Productivity.” hbr.org
- MIT Sloan Management Review & BCG, “Winning With AI.” sloanreview.mit.edu
- McKinsey & Company, “The State of AI” (2025 global survey). mckinsey.com
- Deloitte, “Business and IT Leaders Report AI Agents Are Scaling Faster Than Their Guardrails.” deloitte.com
- MIT Sloan Management Review, “The Three Obstacles Slowing Responsible AI.” sloanreview.mit.edu
- Gartner, agentic AI abandonment forecast, as reported by MarTech. martech.org
- Michael Hammer, “Don’t Automate, Obliterate,” Harvard Business Review, 1990. hbr.org
- Berkeley Artificial Intelligence Research (BAIR), “The Shift from Models to Compound AI Systems.” bair.berkeley.edu
- Databricks, 2024 State of Data + AI Report. databricks.com
- NIST, AI Risk Management Framework Playbook (Govern function). airc.nist.gov
- Joseph Wallace, “The Real Question to Ask About AI Governance,” MIT Sloan Management Review. sloanreview.mit.edu
- Stanford HAI, The 2026 AI Index Report (Responsible AI chapter). hai.stanford.edu
- Boston Consulting Group, “AI Adoption in 2024: 74% of Companies Struggle to Achieve and Scale Value.” bcg.com