Skip to main content
Every Modulate model is audio-native. Each one analyzes the acoustic signal rather than working from transcribed words alone, which is what lets them report tone, accent, synthetic speech, and music alongside what was said.

Velma Triage

Velma Triage: conversation analysis

Detect behaviors, classify conversations, identify participant roles, and extract topics with sentiment. Available as a single API call or a real-time stream.

Models

Models are grouped by the kind of output they produce. Which API should I use? covers the choice between endpoints in detail, including the cases where one call covers two needs.

Transcription

Transcription

Three models across batch and streaming. Speaker diarization, per-utterance timing, and optional emotion, accent, deepfake, and PII/PHI signals.

Detection

Deepfake Detection

Per-frame synthetic-voice verdicts on files or live audio.

Emotion Detection

A whole-file emotion label plus a per-window time series.

Accent Detection

A whole-file accent label plus a per-window time series.

Music & Speech Detection

Frame-level music and speech probabilities.

AI Music Detection

Whether a track contains AI-generated vocals or instrumentals.

Language Detection

The spoken language of a clip, with a confidence score, across 100 languages.

Audio Event Detection

A probability for each of 42 non-speech sound events, from instruments to gunshots.

Redaction

PII/PHI Redaction

A redacted transcript plus audio with the sensitive ranges silenced. Batch and streaming.

New here?

Quick start

Make your first API call in under five minutes. No SDK required.

What you can build

  • Meeting transcription. Multilingual transcripts with speaker labels, timestamps, and optional emotion or accent signals.
  • Live captions. Stream audio over WebSocket and render utterances as they are spoken.
  • Voice agents. Sub-two-second partial transcripts with utterance segmentation at pauses.
  • Anti-spoofing. Real-time deepfake verdicts during a voice authentication flow.
  • Compliance archives. Shareable recordings with PII/PHI removed from the transcript and silenced in the audio.
  • Call QA and coaching. Behavior detections, topics, sentiment, and a summary across a whole conversation.
  • Content moderation. Frame-by-frame music and speech classification at scale.
  • Catalog screening. Batch checks for AI-generated music or AI-generated voice in uploaded audio.
  • Language routing. Identify the spoken language and send audio to the matching pipeline.