Companies that have experienced AI failures in production are paradoxically accelerating their shift toward removing human oversight from deployment decisions, despite these very failures demonstrating the critical value of human judgment. VB Pulse research reveals that 85% of enterprises burned by AI agents that passed evaluations but failed in live environments are now moving faster—not slower—toward automation of go/no-go decisions. This represents a fundamental misalignment between empirical evidence and strategic response: the data shows that evaluation frameworks are failing to predict real-world performance, yet organisations are doubling down on the systems that created the problem in the first place. For CX teams already managing agent deployments through platforms like Zendesk or Salesforce, this trend signals a widening gap between what leadership believes about AI readiness and what actually happens when those systems encounter genuine customer interactions at scale.
The implications cut directly to operational risk in customer-facing environments. If enterprises are removing human gatekeepers precisely when their confidence in automated evaluation should be lowest, support teams will inherit the consequences: more frequent agent failures, higher escalation volumes, and degraded customer experience metrics. The pattern suggests organisations are conflating speed-to-deployment with competitive advantage, treating human review as a cost centre to eliminate rather than a quality control mechanism. This becomes especially acute for mid-market CX operations running multiple AI agents simultaneously—teams lack the engineering resources to build robust evaluation frameworks, yet face pressure to match enterprise deployment velocity. The question becomes whether your organisation's AI governance structure is designed to catch failures before customers do, or whether you're operating under the assumption that your evaluation metrics are sufficiently predictive when the market evidence suggests otherwise.
Human oversight in CX deployment isn't merely a safeguard; it's currently the only reliable validation layer between lab performance and production reality. As organisations strip away these checkpoints, support leaders should be asking whether their current staffing models account for the increased triage burden that will inevitably follow, and whether the cost savings from removing human review actually offset the customer acquisition and retention costs of preventable AI failures.
Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.In July, 13% of 108 enter
Enterprise AI teams have stopped betting on a single orchestration platform. The median enterprise now runs three at once — not by accident, but because none of them fully trusts a single vendor to run the show, according to VB Pulse data.This is not just to avoid vendor lock-in and retain flexibili