TECH2030
Insights
Tech Trends · 5 min readAI-assisted

Why Most AI Pilots Never Reach Production

AI spending is projected to grow 47% this year, surpassing $2.5 trillion, yet only 13% of organizations running more than 100 AI pilots have successfully deployed them enterprise-wide. The root cause lies not in the technology, but in how leadership exercises control.

Key Takeaways

  • A third of organizations have run over 100 pilots, but only 13% reach deployment
  • Failures stem from either overestimating or underestimating AI—ultimately, it's a control problem
  • Goals must be redefined as measurable metrics, not vague directives
  • Even auto-approved outputs need sample audits to confirm actual error rates

AI spending is projected to grow 47% this year, surpassing $2.5 trillion. Yet according to a Zapier survey, nearly a third of organizations that ran more than 100 AI pilots have deployed them broadly across the enterprise in only 13% of cases. About 40% reported that their longest-running AI pilot has remained stuck in testing for over a year.

Jason Gu, Managing Director of AI and Decision Intelligence at BRG, points to leadership control—not the technology itself—as the reason pilots stall, drawing on his own consulting experience. He explains that failures stem from either overestimating AI's capabilities ('AI can do everything') or underestimating the real issues ('the process and data problems we face are due to the model's limitations'), and breaks this down into four leadership control points.

Four Control Points

The first is the starting phase. Vague directives like 'let's use AI to review invoices' rarely go far. Instead, the problem should be redefined around specific business outcomes—such as 'reduce processing time by 95%, achieve 99% accuracy, and route the remaining 1% of exceptions to manual review'—with a single leader accountable for that goal. He emphasized that this accountability should rest with a business leader who has direct experience in that function, not a dedicated AI leader.

The second is pilot design. In the invoice example, the ideal structure involves extracting and normalizing data, then using deterministic logic to cross-reference it against purchase orders and receiving records, automatically approving items that pass according to policy, and routing only exceptions to humans. The choice of document extraction technology—optical character recognition (OCR), multimodal models, or large language models—depends on how variable the documents are. He noted that passing outputs should not automatically be assumed accurate; a meaningful percentage of auto-approved outputs should be audited to determine the actual error rate. He added that common design mistakes include using generative AI where it isn't needed, or building workflows where every single output requires human re-review.

The third is evaluating production economics. Judging cost by price per token or cost of the first response is a mistake. Instead, the benchmark should be the 'fully loaded cost per successful task,' which includes model, tooling, and infrastructure costs, as well as all retries and rework, human review and exception handling, technical integration, evaluation and operations, governance, and the expected losses from incorrect outputs.

Implications for Korean Enterprises Pursuing AX

The metrics highlighted in this piece offer a practical benchmark for Korean companies planning their AX (AI transformation) roadmaps. Running numerous pilots is not, in itself, an achievement—the fact that organizations running over 100 pilots only reached actual deployment 13% of the time shows that what matters isn't the number of pilots, but the design quality of each individual one.

In particular, the observation that failure is often baked in from the outset—when goals are set vaguely, such as 'let's automate this task with AI'—deserves close attention. Companies considering adoption should first redefine their goals as measurable metrics like processing time, accuracy, and exception rates, and then designate a single business leader accountable for those outcomes. Furthermore, scaling up adoption without a process for sample-auditing auto-approved outputs to verify actual error rates risks falling into the same trap as judging costs by token price alone while overlooking total cost of ownership.

Source: 4 leadership pain points that stall AI pilots — and how to fix them

If you found this helpful, share it.

Related insights

More in this category
Enterprise AX
Want to apply this to your own operations?
Request an AX consultation