Inspiration
We came in for the NSA HEARSAY challenge: tell real speech apart from AI-generated speech. Then we made voice clones of ourselves and played them back. They sounded like us. That's what voice-clone scams already use: a call that sounds like your kid or your parent, asking for money right now.
A model that outputs a number doesn't help someone on that call. So we built HocusPocus, a little desktop wizard you can ask to check a voice while you're still talking to it.
What it does
You can drop an audio file on the wizard, or ask it to check a call. If you've given permission, it records 12 seconds of the other person's audio, sends it to our server, and deletes it once it's scored.
You get back likely synthetic, inconclusive, or likely real, along with the likelihood.
On a call, the wizard sits on the call window. Likely real shows a green outline. Likely synthetic shows a purple warning and lets you listen again, message someone you trust, or hang up. If you hang up, the wizard zaps the call window and it shatters, because we built a wizard and had to.
There's also a learning mode with the basics for handling a suspicious call: slow down, call back on a number you know, and set up a family safe word.
How we built it
- App: Electron, with native Swift helpers on macOS that find the call app and capture only its audio output, not your microphone.
- Server: FastAPI on Vultr. It runs the same Python pipeline as our NSA submission and our offline Docker image, so the demo uses the real model.
- Detector: XLS-R 300M, fine-tuned end to end on a mix of real and synthetic speech from a lot of different sources. It scores up to three 4-second windows of 16 kHz mono audio.
- Evidence: Prosody, spectral, voice-quality, and rhythm analyzers show up in the report to explain the result.
- Demo rig: A phone controls a second laptop that places a real call and plays our clips into it through BlackHole and OBS. We tested on actual call audio, not just clean files.
Challenges we ran into
Our interim NSA score was minDCF 0.258, worse than we expected. At that point we were fusing the neural detector with five hand-built forensic analyzers. When we broke down each analyzer's contribution, we found that on hard voice clones the network was pushing toward synthetic and the hand-built analyzers were pulling it back toward real. They had picked up differences between datasets and microphones instead of differences between people and AI. So we changed the final system: the network sets the score, and the analyzers explain it.
Call audio was hard too. Compression, noise removal, and voice isolation all change what the detector listens for. We trained with codec, noise, and time-stretch augmentation and tested through real calls.
The macOS side took a lot of work: detecting calls, capturing one app's audio, following the call window around, permissions, and hanging up reliably.
Accomplishments that we're proud of
- On held-out data (3,546 synthetic and 1,364 real clips, from sources and generators left out of training): AUC 0.9966, EER 2.86%.
- On instant clones of our own voices, the hardest case we tested: minDCF went from 0.48 to 0.18.
- The whole thing works on a live call: detect the call, capture audio with consent, score it with our NSA pipeline, and show the user what to do next.
- The report shows a calibrated likelihood, an inconclusive band, and a warning when call quality makes the result less reliable.
What we learned
One well-trained model on diverse data did better than stacking hand-built techniques.
A lot of what looks like an AI artifact is really a recording artifact. We tested "AI voices don't breathe," and our clones breathed as often as we did (AUC 0.50). We tested the other hints the same way and kept only the ones that held up.
What's next for HocusPocus
- Run the model on the device so checks don't need our server
- Windows and phones, since that's where most scam calls happen
- Better handling of heavily enhanced and noise-removed audio
- Per-window results in the report
- More languages and call conditions
We want HocusPocus to be the thing that makes you stop and call back before you send the money.
Log in or sign up for Devpost to join the conversation.