The real limit of manual voice-agent QA is what teams cannot test at all.
@MavenAGI built its own simulations and smoke tests. Hamming turned them into release infrastructure for concurrency, noise, workflows, and regressions.
“It’s not just time saving, it’s the capability and
We break AI voice agents before your customers do.
Red-teaming • Pre-launch testing • Production monitoring
- Voice agents are built to help. That is also what makes them exploitable. Give an agent access to customer data and permission to act, and a normal conversation becomes an attack surface. Watch @sumanyu explain why teams should red-team voice agents:
- A production voice agent needs more than a model. @OpenAI Presence shows what sits around it: policies, actions, evals, escalation rules, production signals, and controlled rollouts. Our addition: turn failed production calls into regression tests before the next release.
- A voice agent can score 94% in the lab and 58% in a moving car. Only the background noise changed. Across 10M+ call minutes, ASR accuracy drops 20-40% once noise passes 10dB SNR. If you don't know your test calls' SNR, you don't know how the agent behaves in the real world. 👇
- Compliance dashboards look great until the auditor asks a question they were never designed to answer. Which policy version? Who reviewed it? Was it fixed? Did it come back? We wrote a guide for building evidence packets that hold up. 👇

