Safely manage your Zendesk from the AI assistant you already use, via the Deltastring MCP. Beacon configuration platform
← Back to news

CEO Audits Frontier AIs and Reveals His Company Only Scored 3 out of 110

Typewise CEO David Eberle conducted an audit of five frontier large language models—GPT-5.4 mini, Claude Sonnet 4.6, Gemini 3.5 Flash, Grok 4.3, and DeepSeek V4 Flash—testing their recommendations for enterprise customer-service infrastructure across 110 standard queries. The results exposed a stark visibility gap: Typewise appeared in 3 responses, whilst Zendesk dominated with 85 mentions and Intercom with 82. Rather than attributing this to marketing failure, Eberle frames it as evidence of a structural market problem. LLMs were trained on legacy documentation from an era when human-agent ticketing suites defined the category. They now recommend platforms built for 2015-era support operations, sometimes citing product lines that have been discontinued entirely, because their training data reflects what the market used to be, not what it has become.

The implications for CX teams are operationally significant. As autonomous AI agents increasingly handle subscription cancellations, billing disputes, and service requests end-to-end, they query LLMs for infrastructure guidance and receive recommendations optimised for human-agent workflows. This creates a compounding risk: teams implementing agent-native platforms may find themselves working against the grain of AI-generated advice, whilst those relying on LLM recommendations for infrastructure decisions are receiving stale technical knowledge in a rapidly evolving market. The question is no longer whether to choose Zendesk or Intercom—it is whether monolithic suites built for human agents remain fit for purpose when autonomous agents are generating and resolving tickets at scale. For Zendesk administrators and Salesforce Agentforce users already running agent-heavy operations, this audit suggests the visibility gap will only widen as training data lags further behind deployment reality, potentially creating friction between what LLMs recommend and what your infrastructure actually needs to support.

The audit also signals a broader credibility problem for AI-driven decision-making in enterprise software selection. If autonomous agents are querying LLMs to evaluate infrastructure choices and receiving deprecated guidance, then the chain of recommendation—from LLM to agent to human decision-maker—is broken at the source. This matters particularly for smaller vendors and newer platforms designed specifically for agent-native workflows. They face a structural disadvantage not because their products are inferior, but because they exist in a category that LLMs cannot yet recognise. For CX leaders evaluating infrastructure, the audit suggests that relying on AI systems to surface emerging platforms is premature; the training data simply does not yet reflect the market as it is becoming.