Two tools for evaluating a voice agent

In its 5 October 2026 update, Vapi added simulation latency limits for the model, voice and turn timing. Exceeding a limit fails the test, and results show measured latency. The Soniox transcriber also gains a configurable confidence threshold for discarding uncertain transcripts, available through the interface and API.

VoiceIP analysis: test unclear answers deliberately

An online store could prepare examples where a customer gives an order number against background noise, corrects a date or pauses for a long time. Define the expected response in advance: ask for clarification, repeat a detail or hand the call to a manager. Discarding uncertain text does not itself define that behaviour.

Include clear, brief answers too. An overly strict threshold could cause unnecessary clarification requests in your scenario. Compare settings against the same examples and retain the results so that a subsequent voice or model change does not pass unnoticed.

What to measure before real calls

Record separate expectations for the pause after a customer speaks and the wait for CRM data. Include cases where the order service is slow or unavailable. Check the final order status alongside response speed. A successful simulation is only one stage of evaluation: controlled phone calls are still needed before launch. This is an editorial testing plan, not a promise of a particular outcome.

Sources