Safely manage your Zendesk from the AI assistant you already use, via the Deltastring MCP. Beacon configuration platform
← Back to news

AI chatbots are taking over customer service. But how often are they wrong?

AI chatbots are now the primary customer service channel for most businesses, yet independent research indicates roughly half of all responses contain errors, misleading information, or missing context. The NBC 5 investigation documented a concrete example: United Airlines' chatbot told customer Alison Gil her $200 travel credit would expire in five years, when the actual policy allowed only one year. Gil discovered the discrepancy weeks later via automated email reminder, leaving her with months rather than years to use the credit. When she contacted United, the airline offered no explanation for the chatbot's error and refused to extend the expiration date, citing policy constraints. The incident reflects a broader pattern visible in online forums, where customers report chatbots providing information so inaccurate it triggered unintended consequences—one case involved a refund inquiry that resulted in an entire trip cancellation.

The accountability gap here is the critical issue for CX teams. V.S. Subrahmanian, a computer science professor at Northwestern, notes that whilst human operators and AI systems both make mistakes, organisations must accept responsibility for those errors. Chatbots fundamentally lack the dynamic listening capability humans possess—the ability to sense when information sounds wrong and course-correct in real time. United's response to the NBC investigation revealed the core problem: the airline acknowledged the chatbot provided incorrect information but offered only a generic disclaimer that "information may not always be complete or relevant." This raises a pressing question for teams already operating AI-driven support: if your chatbot is the sole customer service option, how are you measuring accuracy rates, and what escalation protocols exist when errors occur? The current landscape suggests most organisations are deploying these systems without adequate guardrails, relying instead on reactive damage control when customers surface problems.

The regulatory and operational implications are substantial. Consumer advocates now recommend customers screenshot chatbot conversations as proof in case information proves inaccurate—a workaround that essentially shifts verification burden to customers rather than organisations. For CX professionals managing platforms like Zendesk or Salesforce Service Cloud, this signals that AI automation without human oversight creates liability rather than efficiency gains. The question becomes whether your team's implementation includes mandatory human review for high-stakes interactions, knowledge base validation against actual policies, and clear escalation pathways when chatbot confidence scores fall below acceptable thresholds. Without these controls, you're operating a system that half the time delivers incorrect guidance—a ratio that would be unacceptable for human agents but somehow remains normalised for AI.