Google's release of Gemini 3.6 Flash and its companion models represents a direct assault on the operational cost structure of AI-powered customer service. The 65% reduction in token consumption for long-horizon tasks—the exact workload profile that defines complex support interactions—fundamentally reshapes the economics of deploying agentic AI in contact centres. At $1.50 per million input tokens, this pricing tier makes continuous agent operation economically viable at scale for organisations that previously faced prohibitive per-interaction costs. For CX teams already running Agentforce or similar enterprise platforms, this creates immediate pressure: the cost-per-resolution advantage that justified premium vendor pricing erodes when commodity models deliver comparable performance at a fraction of the expense.
The efficiency gains matter less than what they enable operationally. Token efficiency directly translates to faster response times and reduced latency in multi-turn conversations—the hallmark of effective support interactions where context accumulates across exchanges. This is particularly acute for technical support, billing disputes, and policy clarification scenarios where agents must maintain coherent reasoning across 10+ conversational turns. The question for support leaders becomes whether their current vendor partnerships can match this cost trajectory, or whether the economics now favour building custom agent layers atop cheaper foundational models. Smaller CX platforms that lack proprietary model development capacity face genuine competitive pressure here.
The broader implication sits in operational flexibility. Teams no longer face a binary choice between expensive, capable models and cheap, limited ones. Gemini 3.6 Flash's performance on engineering tasks—proxy for complex problem-solving—suggests CX organisations can now run sophisticated agents continuously without the budget gymnastics that previously required careful token budgeting or interaction gating. This shifts investment priorities from model licensing toward orchestration, routing logic, and integration quality—areas where differentiation actually matters for customer outcomes rather than raw inference cost.
Google DeepMind today released three new proprietary AI models it says are among its most token-efficient yet: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The models aim to make AI agents faster, smarter, and cheaper at scale. Google is pricing Gemini 3.6 Flash at $1.50 per