Safely manage your Zendesk from the AI assistant you already use, via the Deltastring MCP. Beacon configuration platform
← Back to news

Zendesk Thinks AI Customer Service Has a Measurement Crisis

Zendesk

Zendesk's announcement of Quality Score exposes a fundamental disconnect in how contact centers evaluate AI-driven customer support: they're measuring the wrong things at the wrong scale. Forrester's finding that Net Promoter Score declined across 20 of 39 industry-country combinations in 2025 despite 92% of contact centers operating quality assurance programs reveals the core problem. Traditional QA systems were built around sampling 2–5% of interactions when human agents handled manageable volumes. AI agents process thousands of conversations daily, rendering a 2% sample statistically meaningless—effectively 0.001% visibility into actual performance. The issue isn't the absence of measurement frameworks; it's that teams have optimised for metrics that look clean on spreadsheets (response time, deflection rates, cost per interaction) rather than outcomes that matter to customers (problem resolution, human-like experience, satisfaction). This measurement gap became visible through Klarna's public stumble, where an AI chatbot replaced 700 agents across 23 markets in January 2024, yet the headline metrics masked deeper quality issues.

What separates observation from action is the feedback loop itself. Dave Giblin's framing—that "measurement alone is observability" and only the continuous cycle between signal and fix drives outcomes—positions Quality Score as infrastructure for closed-loop improvement rather than another dashboard metric. For teams already running Agentforce or similar agentic platforms, this raises an uncomfortable question: are your current QA programs actually equipped to surface the patterns that matter at AI scale, or are they simply validating decisions already made? The implication for CX leaders is stark. Full-conversation measurement becomes mandatory, not optional, but only if it's paired with operational discipline to act on what the data reveals. Organisations that treat Quality Score as a compliance checkbox will replicate Klarna's problem at different scales. Those that embed it into continuous coaching and model refinement cycles will gain genuine competitive advantage as customer expectations continue to outpace incremental service improvements.