Production QA checklist
Build a QA loop that leads from a bad call to a verified fix.
Use production conversations, consistent evaluation criteria, evidence-led review, and follow-up checks to improve a voice agent without relying on random call sampling.
Before you evaluate
- Define the customer outcome for each supported workflow.
- Choose representative calls, including failures and edge cases.
- Confirm provider permissions, retention, and evaluator data-handling choices.
- Document what pass, fail, and uncertain mean for the scorecard.
While you review
- Use scores to prioritize calls, then inspect transcript and available audio together when timing matters.
- Require evaluator rationales to point to specific conversation evidence.
- Escalate uncertain, sensitive, or consequential outcomes to a human reviewer.
- Separate agent behavior from provider or infrastructure failures.
After an agent change
- Group recurring failures by likely root cause and assign corrective work.
- Retest against relevant historical examples and new production calls.
- Compare quality signals over time, not only operational activity.
- Keep the calls behind a regression visible to the people making the next change.
Own your benchmark
Choose voice models using evidence from your own calls.
See how a private benchmark runs inside your environment, on your production scenarios, across the models you are considering.
30 minutes. Bring your model-selection question.