Safely manage your Zendesk from the AI assistant you already use, via the Deltastring MCP. Beacon configuration platform →
← Back to news
ai

Agentic reliability and evaluations : Enterprises that got burned by a bad eval are the most likely to remove humans from the loop, not the least

Enterprises are doubling down on automation despite mounting evidence that their evaluation frameworks are broken. Across 108 organizations, trust in automated agent evaluation nearly tripled between June and July—from 5% to 13%—whilst the actual failure rates these evaluations purport to measure remained flat. This disconnect reveals a critical pattern: teams that have experienced failures in their agentic systems are not becoming more cautious about removing human oversight. Instead, they're accelerating automation adoption, likely because they've invested heavily in agent infrastructure and see human-in-the-loop processes as bottlenecks rather than safeguards. The paradox is stark: the very enterprises that should be most sceptical of their evaluation capabilities are the ones most aggressively eliminating human judgment from their workflows.

For CX leaders managing Zendesk, Salesforce Service Cloud, or similar platforms, this trend carries immediate operational risk. If your team has recently deployed agentic features—whether through Agentforce, Freshdesk's AI capabilities, or custom integrations—you're likely experiencing the same evaluation blind spot affecting these 108 enterprises. The question becomes whether your current metrics are actually measuring customer outcomes or merely measuring what your evaluation system thinks it's measuring. Teams that have suffered agent failures are particularly vulnerable to this trap: the pressure to prove ROI and justify the initial investment often outweighs the discomfort of admitting that evaluation frameworks need rebuilding. This creates a dangerous feedback loop where confidence in automation rises precisely when it should be questioned most rigorously.

The implications extend beyond individual team performance. If enterprises are systematically overestimating their ability to evaluate agent reliability, then the broader agentic ecosystem is operating on false confidence. This affects vendor selection, feature prioritization, and ultimately customer experience quality across the industry. CX teams should treat this moment as a forcing function: audit your evaluation criteria now, before the gap between perceived and actual agent performance becomes a compliance or reputational liability. The enterprises getting burned aren't the ones pulling back—they're the ones pushing harder, which means the real risk isn't in adoption, but in adoption without adequate measurement discipline.