Enterprises are doubling down on automation despite mounting evidence that their evaluation frameworks are broken. Across 108 organizations, trust in automated agent evaluation nearly tripled between June and July—from 5% to 13%—whilst the actual failure rates these evaluations purport to measure remained flat. This disconnect reveals a critical pattern: teams that have experienced failures in their agentic systems are not becoming more cautious about removing human oversight. Instead, they're accelerating automation adoption, likely because they've invested heavily in agent infrastructure and see human-in-the-loop processes as bottlenecks rather than safeguards. The paradox is stark: the very enterprises that should be most sceptical of their evaluation capabilities are the ones most aggressively eliminating human judgment from their workflows.
For CX leaders managing Zendesk, Salesforce Service Cloud, or similar platforms, this trend carries immediate operational risk. If your team has recently deployed agentic features—whether through Agentforce, Freshdesk's AI capabilities, or custom integrations—you're likely experiencing the same evaluation blind spot affecting these 108 enterprises. The question becomes whether your current metrics are actually measuring customer outcomes or merely measuring what your evaluation system thinks it's measuring. Teams that have suffered agent failures are particularly vulnerable to this trap: the pressure to prove ROI and justify the initial investment often outweighs the discomfort of admitting that evaluation frameworks need rebuilding. This creates a dangerous feedback loop where confidence in automation rises precisely when it should be questioned most rigorously.
The implications extend beyond individual team performance. If enterprises are systematically overestimating their ability to evaluate agent reliability, then the broader agentic ecosystem is operating on false confidence. This affects vendor selection, feature prioritization, and ultimately customer experience quality across the industry. CX teams should treat this moment as a forcing function: audit your evaluation criteria now, before the gap between perceived and actual agent performance becomes a compliance or reputational liability. The enterprises getting burned aren't the ones pulling back—they're the ones pushing harder, which means the real risk isn't in adoption, but in adoption without adequate measurement discipline.
Across 108 enterprises, trust in automated agent evaluation rose sharply in July — and the failure rate it is supposed to predict did not move at all. The share of organizations that fully trust automated evaluation nearly tripled, from 5% in June to 13%, and the complaint that evaluations don’t mat