Safely manage your Zendesk from the AI assistant you already use, via the Deltastring MCP. Beacon configuration platform
← Back to news
ai

New agentic memory framework uses 118K tokens per query. LangMem burns through 3.26M.

A new memory framework developed at the National University of Singapore exposes a critical inefficiency in how current AI agents handle customer service workloads. Existing systems like LangMem consume 3.26 million tokens per query by relying on static retrieval pipelines that flood context windows with noise, forcing agents to sift through irrelevant information before reasoning through customer problems. MRAgent, the alternative framework, reduces this to 118K tokens per query by abandoning the traditional "retrieve-then-reason" approach in favour of dynamic, iterative memory management. For CX teams already invested in agentic platforms—whether Salesforce's productized customer service agent or RingCentral's expanded agentic capabilities in RingCX—this represents both a technical problem and a cost problem that will soon become impossible to ignore.

The implications cut directly to operational economics and agent performance. Token consumption at LangMem's scale translates to substantial API costs, slower response times, and degraded reasoning quality as agents struggle to extract signal from oversized context windows. A 27x reduction in token usage per query fundamentally changes the unit economics of agentic AI in customer service, making it viable for high-volume contact centres that currently treat AI agents as experimental pilots rather than production infrastructure. The question facing support leaders is whether their current vendor roadmaps account for this efficiency gap—and whether platforms that haven't optimised memory retrieval will become cost-prohibitive as query volumes scale.

This efficiency breakthrough also reshapes vendor positioning in the agentic AI space. Platforms that can demonstrate token efficiency comparable to MRAgent will capture teams evaluating long-term deployment costs, whilst those relying on brute-force retrieval will face pressure to either innovate or justify premium pricing. For mid-market CX operations considering whether to build custom agents or adopt vendor solutions, the token consumption profile should now be a primary evaluation criterion alongside accuracy and latency.