Zendesk experienced two distinct but operationally significant incidents within 24 hours in mid-September 2026. The first, affecting Support customers across multiple pods on September 16, rendered custom ticket statuses inactive and irretrievable for over three hours, crippling agents' ability to progress tickets through standard workflows and execute macros. A production code change inadvertently removed a required setting that determined feature availability, forcing the platform to treat custom statuses as disabled. The second incident, isolated to Pod 20 Voice customers on September 17, trapped agents in wrap-up mode following calls, preventing them from accepting new inbound work and blocking call recordings from appending to tickets. This stemmed from a single account generating abnormally high volumes of duplicate agent availability updates that overwhelmed shared processing capacity on the shard cluster. Both incidents were resolved through rollback and temporary feature disabling respectively, but their root causes expose systemic vulnerabilities in Zendesk's deployment and resource management architecture.
The implications for CX teams are twofold. First, these incidents demonstrate that foundational features—ticket status management and agent availability—remain susceptible to cascading failures despite their criticality to daily operations. For teams already dependent on custom status workflows or those managing high-volume voice operations, the question becomes whether current SLA commitments adequately account for the blast radius of such outages. A three-hour disruption to ticket statuses across multiple pods affects not just individual agents but entire escalation and routing logic; similarly, agents stuck in wrap-up mode directly translate to abandoned inbound calls and degraded service levels. Second, Zendesk's remediation roadmaps reveal concerning gaps in pre-deployment validation and rate-limiting safeguards. The Support incident required adding checks to confirm features continue working post-update—a control that should exist before production deployment. The Voice incident exposed the absence of per-user update frequency limits and automated pause mechanisms for abnormal activity patterns. For support leaders evaluating platform stability or planning capacity, these incidents suggest that Zendesk's infrastructure, whilst generally resilient, lacks sufficient guardrails against both human error in deployments and resource exhaustion from single-account anomalies.
The timing and nature of these failures carry broader significance for the CX vendor landscape. Zendesk's multi-pod architecture is designed to isolate failures, yet the Support incident affected multiple pods simultaneously, indicating the issue propagated across infrastructure boundaries. Conversely, the Voice incident remained isolated to Pod 20, suggesting segmentation worked as intended—but only after the fact. For mid-market and enterprise teams considering platform consolidation or migration, these incidents underscore the importance of understanding how vendors handle cross-pod deployments and whether rate-limiting and anomaly detection are baked into core services or bolted on reactively. The remediation items published by Zendesk are notably defensive rather than innovative, focusing on preventing recurrence rather than improving resilience. This raises a critical question for CX leaders: should platform stability and incident response speed be weighted more heavily than feature velocity when evaluating vendor roadmaps?
SummaryBetween 14:26 UTC and 17:40 UTC on 17 September 2026, Voice customers on Pod 20 experienced an issue that caused some agents to remain in wrap-up mode after calls. Affected agents were not offered new incoming calls, which could delay inbound call handling.TimelineSeptember 17, 2026 16:53 UTC
SummaryOn September 16, 2026, from 09:00 UTC to 12:20 UTC, Support customers on multiple pods experienced an issue that caused custom ticket statuses to appear inactive or become unavailable. As a result, agents were unable to use or reactivate these statuses, disrupting routine ticket updates, macr