Safely manage your Zendesk from the AI assistant you already use, via the Deltastring MCP. Beacon configuration platform
← Back to news

Service Incident - July 16, 2026 - Chat | Pod 25

Zendesk

A single customer's high-volume ticket update automation caused a 106-minute outage affecting Chat ticket creation across Pod 25 on July 16, 2026, from 03:50 to 05:36 UTC. The incident stemmed from unexpected database load on the Ticket Log Service when one account's automation generated a sudden surge in ticket updates, creating cascading latency that degraded the Agent Workspace experience for all customers sharing that infrastructure pod. Zendesk's response was swift—engineers implemented temporary rate limiting within 46 minutes of the initial alert, restoring normal performance and preventing further degradation. Notably, the Messaging product remained unaffected, suggesting adequate isolation between product lines, though the incident exposes a critical vulnerability in pod-level resource sharing.

The root cause reveals a systemic weakness in multi-tenant architecture: the absence of sufficient safeguards to prevent a single customer's runaway automation from impacting neighbours on the same pod. This raises an uncomfortable question for teams already operating at scale—how many of your own automations might be generating similar load patterns without triggering alerts? Zendesk's remediation roadmap addresses this directly through improved query-level monitoring, query optimisation, and targeted throttling mechanisms. However, the reliance on temporary rate limiting as an immediate fix suggests these protections were not in place beforehand, a gap that should concern any organisation running mission-critical support operations on shared infrastructure.

For CX leaders, this incident underscores the operational risk of pod-based architecture during periods of automation growth. As teams increasingly automate ticket workflows—particularly through integrations and custom automation—the likelihood of triggering similar load spikes increases proportionally. The incident also highlights why visibility into your own automation patterns matters: teams should audit their ticket creation and update workflows to understand their own load profiles before Zendesk's preventive measures are fully deployed. Until query-level monitoring and throttling are universally implemented, the burden of preventing such incidents falls partly on customers to self-regulate their automation intensity.