A consistent pattern has emerged in the evolution of voice automation: that while speech recognition and conversational interfaces have improved significantly, the underlying operating model for automation has remained largely unchanged.
To find out more about the challenges this is posing, we asked independent researcher and practitioner Ankit Talwar to put the spotlight on why most voice automation systems appear to be falling short right now and what should be happening instead.

Most systems still optimize for routing efficiency rather than resolution quality. This structural limitation becomes most visible in one of the most persistent challenges spanning contact centres and supply chains alike: Where Is My Order (WISMO).
Across retail and e-commerce, delivery-status enquiries consistently represent a significant share of inbound customer contacts and rise sharply during peak demand periods.
Industry research indicates that WISMO enquiries can account for 25–35% of total contact volume, with spikes exceeding 50% during seasonal peaks.[1][2]
Despite decades of automation investment, these contacts remain resistant to meaningful containment. The issue is not the maturity of voice technology; it is the architecture around it. Routing-first automation models are the primary structural failure mode preventing progress.
WISMO is often treated as a simple informational request. Customers want to know when their order will arrive. In practice, this framing significantly underestimates the problem.
Definition: This article defines WISMO not as a call type, but as an operational signal of upstream data failure.
Customers are rarely calling to hear a date repeated. They are calling to decide whether the delivery still works for them and, if not, whether alternatives exist. When automation can only restate status without supporting a decision, escalation becomes inevitable.
Three structural forces drive cost and friction in WISMO interactions:
Most voice automation systems rely on order-centric platforms to answer delivery questions. These systems perform well at answering when an order is expected to arrive.
Where they consistently fall short is explaining why an outcome changed and what can be done about it.
When deliveries stall, dates shift, or tracking information stops updating, the explanation typically resides outside the order record itself.
It lives within broader operational systems that reflect execution reality rather than planned intent. Order-centric automation can predict outcomes; it cannot support decisions.
To find out why the relationship between personalization and automation matters, read our article: Balance Automation and Personalization in CX
The information customers need to make decisions resides in operational signals such as:
Without access to this operational context, automation can only repeat a status. It cannot determine whether a delivery can be rescheduled or if a “hold-at-location” option exists.
As a result, customers escalate to human agents, who must reconstruct the same context manually. This reconstruction effort is where time, cost, and customer trust erode.
The Context-First Resolution Framework is a decision architecture that shifts automation away from routing efficiency and toward decision readiness.
Within this framework, Voice AI handles high-volume, time-sensitive enquiries, identifying intent and retrieving relevant operational context. Critically, Voice AI stops when a decision or trade-off is required.
At the centre of the framework is a Multimodal Context Engine. This functions as a persistent layer that preserves the intersection of customer intent and operational reality. Unlike traditional omnichannel set-ups that reset context at hand-off, this approach ensures continuity.
Industry research shows that organizations are increasingly redesigning service journeys so that context flows seamlessly between AI systems and human agents.[5] The result is not faster routing; it is better decisions.

Here is an illustrative decision pattern for delayed fulfilment; a delayed delivery caused by a missed pick-up:
When escalation occurs, the agent receives decision-ready context rather than raw data. This reframes agent assistance from “retrieval work” to “decision execution.”
Gartner projects that conversational AI deployments could reduce agent labour costs by $80 billion by 2026.[4] However, these gains will not come from deploying voice interfaces alone.
Organizations that fail to close the context gap – the loss of operational truth between systems and people – will see automation plateau at routing rather than resolution.
As enterprise AI matures, the boundary between contact centre operations and supply chain execution will dissolve. WISMO sits precisely at that intersection.
The Context-First Resolution Framework treats WISMO not as a volume problem to be deflected but as a coordination challenge to be solved through unified intelligence. Together, Voice AI, Multimodal Context, and Human Insight transform WISMO from a volume problem into a solvable one.
Written by: Ankit Talwar, Director of Product Management for AI at Dell Technologies and a Distinguished Fellow of the Soft Computing Research Society (SCRS)
For more information on contact centre automation and technology, read these articles next:
Reviewed by: Xander Freeman