Safely manage your Zendesk from the AI assistant you already use, via the Deltastring MCP. Beacon configuration platform
← Back to news

AI Slop Is a Customer Experience Problem – Is Multimodal the Cure

Text-only AI in customer support has created a structural problem that most CX teams haven't yet accounted for: whilst organizations have reduced operational costs and expanded coverage, customers now bear a hidden cognitive load that doesn't register on traditional dashboards. This "effort tax" manifests as customers parsing walls of generated text, verifying instructions before acting on them, and rephrasing questions when the AI misses context. The real damage emerges through "AI slop"—hallucinated instructions that can cause product damage, safety issues, or liability exposure, as evidenced by Woolworths' false personal claims, Air Canada's incorrect refund guidance, and DPD's brand-damaging chatbot failures. These incidents reveal that the structural weakness of language-based AI isn't a prompt engineering problem; it's that language itself presents "a tree of possibilities" where the AI selects one interpretation, and when that choice is wrong, the customer pays the correction cost. For teams already running text-heavy implementations across Zendesk, Freshdesk, or Salesforce Service Cloud, this raises an uncomfortable question: are your containment rates and CSAT improvements masking a deterioration in actual customer trust?

Multimodal support—integrating images, video, and real-time visual feedback—addresses this at a structural level by constraining interpretation through visual grounding. A photograph of a faulty component, a live error display, or a video showing incorrect installation technique doesn't give the AI ambiguity; it provides a scene that's harder to misinterpret. Real-time visual feedback solves the delayed-feedback failure mode where customers follow incorrect text instructions for 10-15 minutes before discovering the error, compounding frustration and eroding confidence in subsequent interactions. The six properties of effective multimodal support—enhanced context, reduced ambiguity, cross-modal consistency, state awareness, real-time feedback, and visual grounding—work together to eliminate hallucination through information diversity rather than better prompting. For support leaders evaluating platform upgrades or AI vendor partnerships, the critical question becomes whether your current tooling can support multimodal workflows at scale, or whether you're locked into text-only architectures that will increasingly underperform as customer expectations shift from "faster resolution" to "trustworthy resolution."