SilowsilowBook a demo
Insights

Why 40% of agentic AI projects will be scrapped — and the boring reason it keeps happening

10 July 2026
A row of verified squares whose ticks stop two thirds of the way across, above a measuring rule that runs out of markings at the same point.

TL;DR: Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027. The usual cause is not the model. Teams automate without a record of how the work is done today, so there is no ground truth to build against or measure by. Before you pick a project, map how work moves through your team and rank the candidates by payback.

The prediction gets quoted at every conference: Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. It lands as a warning about the technology. It isn’t. Underneath all three of those reasons sits a missing step — and the step is unglamorous enough that almost everyone skips it.

Here’s the pattern, and it repeats with unnerving consistency.

The shape of a failure

A team picks a workflow. Someone builds an agent. It demos beautifully — the demo is always beautiful, because the demo is the happy path and the person driving it knows exactly which inputs behave.

It goes to a pilot. In the pilot, roughly four things happen:

  1. It works on the cases the builder imagined.
  2. It fails quietly on the ones they didn’t.
  3. Nobody notices for two weeks, because there’s no definition of “correct” to measure against.
  4. Trust collapses, and it collapses all at once.

Trust is not a gradient. People don’t downgrade an unreliable agent to “use with caution” — they stop using it. And once they’ve stopped, the project is dead regardless of how good the underlying model was.

The agent didn’t fail an evaluation. It was never given one.

The step everyone skips

Every mature engineering discipline has an oracle — some source of truth that says what the right output was supposed to be. Tests have expected values. Compilers have specifications. Even a spreadsheet has a number you can check against reality.

Agentic work shipped without one. We built systems that take fuzzy input, produce fuzzy output, act on the world — and then evaluated them by asking whether the demo felt good.

Why “eval” got skipped

It’s not laziness. It’s that the ground truth was genuinely hard to get:

  • The task lives in someone’s head. The senior person who handles the tricky refunds doesn’t have a written procedure. They have judgment.
  • Correctness is contextual. The right answer for this customer is the wrong answer for that one, and the difference isn’t in the ticket text.
  • The happy path is the smallest part. Most of the value — and all of the risk — lives in the exceptions.

So teams shipped without a benchmark, because building the benchmark looked harder than building the agent. It was harder, and it still is.

Where ground truth actually comes from

Here’s the part that took us a while to see clearly.

If you have watched how the work is done — really watched it, at the level of what happened, in what order, with which decisions — then you already have the benchmark. Not a synthetic test set. A recording of a competent human doing the job correctly, dozens of times, including the exceptions.

That record answers the question the demo can’t: did the agent do what the person would have done?

What that record gives you

A record of the work done well is not a feature of any one tool. It’s the input that makes the rest of the process possible:

  • Situations that actually occurred, rather than the ones the builder imagined — so you can check the agent against reality instead of against a wish.
  • A comparison to what the person did, step by step, so a divergence is something you can point at rather than argue about.
  • A basis to decide before production whether those divergences are acceptable — and to refuse to ship when they aren’t.

None of this requires a new model. It requires having looked at the work first.

The uncomfortable implication

If your evaluation set is a spreadsheet someone wrote from memory on a Tuesday, your agent is graded against a fiction. It will pass. It will still fail in production, and you will conclude that agents don’t work — when what didn’t work was the grading.

The projects that survive won’t be the ones with the best models. Everyone rents the same models. They’ll be the ones that could answer, before shipping, a question most teams still can’t:

What does correct look like, and how do we know?

Start from the work

The order matters more than the tooling. Observe the work. Rank where AI pays back. Build the thing. Validate it against what you observed. Then, and only then, put it in front of people.

Skip the observation and every later step is guesswork wearing a confidence interval.

If you want to see what that looks like on your own workflows — book a call.

Nikita Sorokin
Written by
Nikita Sorokin
Founder & CEO, Silow

Nikita started Silow after watching company after company buy AI it never used — the tools were fine, but nobody could say which work was worth automating. He leads product and the company, and writes most of what you read here.

See how Silow maps your work into a ranked roadmap.

Book a call