Evaluation workspace
Your calls are the only benchmark that matters.
VaaniEval replays your real production scenarios through candidate STT, LLM, and TTS models and shows exactly where they disagree, how much it costs, and what it changes for the customer.
A benchmark built on generic data cannot predict your calls
Public leaderboards are measured on curated datasets. Your traffic is not curated. It carries regional accents, code-switching, domain vocabulary, telephony compression, background noise, interruptions, and the specific details that drive an action: names, amounts, dates, and identifiers.
Private benchmark set
Select representative production calls and freeze them into a reusable evaluation set with the outcomes that mattered.
Identical scenarios
Run the same audio and conversation state through your current stack and every candidate model, offline or in shadow.
Turn-level evidence
Inspect the audio, transcripts, word-level disagreements, tool decisions, timing, and the response the caller would have heard.
What VaaniEval measures
- Transcription quality where it counts. Aggregate error rates plus flagged disagreements on names, numbers, dates, locations, and domain terms.
- Conversational latency. Endpointing behavior, time-to-first-token, time-to-first-audio, and turn latency distribution, not just averages.
- Downstream decisions. Whether a transcript difference changed the LLM response, the tool call, or the final outcome.
- Speech output quality. TTS pronunciation of critical entities, prosody on numbers, and failure modes in your target languages.
- Cost and reliability. Cost per call, error and timeout rates, and provider stability across the same run.
Offline replay and real-time shadow
Offline replay is the cheapest way to compare models across a large historical set. Real-time shadow evaluation runs candidates alongside live traffic so you can observe behavior under real conditions before a customer ever hears the new stack.
Shadow evaluation raises provider spend and, if you select hosted models, creates an external data path. VaaniEval treats that as a governed decision with routing controls, redaction, budgets, and audit logs rather than a hidden default. See how deployment and routing work.
Model-to-model disagreement tells you where stacks diverge and which differences carry risk. For high-stakes accuracy claims, pair it with a human-reviewed reference set. We will say this in the demo too.
Explore the evaluation metrics guideOwn your benchmark
Choose voice models using evidence from your own calls.
See how a private benchmark runs inside your environment, on your production scenarios, across the models you are considering.
30 minutes. Bring your model-selection question.