Modulate’s cover photo
Modulate

Modulate

Software Development

Somerville, Massachusetts 7,557 followers

Frontier AI lab building Velma — audio-native voice intelligence for developers and platforms

About us

Modulate is a frontier AI lab building the voice intelligence layer for the internet. We develop foundation models for audio: models that don't just transcribe speech, but understand it: who's speaking, what's being said, whether it's real, and whether it's safe. Velma is Modulate's audio-native voice intelligence model, purpose-built for audio from the ground up — not a text model retrofitted for speech. Velma powers a suite of APIs for developers building the next generation of voice-first products: → Modulate Transcribe: state-of-the-art speech-to-text. #1 on the Hugging Face Open ASR Leaderboard (out of 88 models), with top results on Sierra's μ-Bench, including #1 for Mandarin (zh-CN) → Deepfake Detect: real-time detection of AI-generated and cloned voices → AI Music Detection: identifying AI-generated audio content at scale → PII/PHI Redaction: automated sensitive-data protection for regulated industries Developers use Velma to build voice products that are faster, more accurate, and safer by default — without stitching together brittle, single-purpose audio tools. ToxMod, our proactive voice moderation system, applies this same model foundation to real-time safety — trusted by studios including Ubisoft, Activision, and Schell Games to detect toxicity and harassment without relying on player reports. Modulate was founded on a simple premise: audio is its own modality, and it deserves models built natively for it - not adapted from something else.

Website
http://modulate.ai
Industry
Software Development
Company size
51-200 employees
Headquarters
Somerville, Massachusetts
Type
Privately Held
Founded
2019

Employees at Modulate

View 50 employees at Modulate

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

See all employees

Locations

  • Primary

    Suite 300, Somerville, MA 02144, United States

    Somerville, Massachusetts 02144, US

    Get directions

Updates

  • Modulate reposted this

    When Tom Elliott came to me and asked VoiceRun to sponsor a Voice AI Meetup series in Boston, it was an obvious yes. Asking Mike Pappas & Modulate to help out too? Another obvious yes. And of course BUILD617 was down to host. Voice is still an underinvested interface across tech, including in Boston. The Voice AI includes voice agents, but it also includes media pathing, QA for human agents, analytics and insights, and more. The impact can in phone calls, but it call also be in apps. There is a ton of untapped opportunity for those brave enough to seize it! Tom is the perfect guy to organize this. He's an immensely creative Voice developer in Boston who is stepping up to bring the connective tissue of Boston together. I'll see you there! Comment/DM if you request access, and I'll put a good word in :). https://luma.com/gqxuajjg

  • Clinical audiology transcription needs more than words on a page - it needs context. One G2 reviewer built an AI Scribe SaaS on Velma 2 for exactly that reason: ⭐ 5/5 "Accurate Transcriptions with Emotional Context at a Fantastic Price" Live diarization. Emotional context per utterance. Accuracy that holds up in regulated, high-stakes use cases. That's Velma 2, built for developers who need transcription that understands more than just what was said 👀 See what you can build 👇

    • No alternative text description for this image
  • Speech-to-speech isn't the only way to build a natural-sounding AI voice agent. That's a common misconception 🚨 The argument for speech-to-speech is understandable: if an AI agent needs to respond naturally, it needs to understand more than *what* a person said. It also needs to understand *how they said it*. Their emotion. Their prosody. Their hesitation. The pauses between words. The things that don't appear in a traditional text transcript. The problem isn't necessarily the cascade architecture. The problem is flattening speech into plain text and throwing that information away. An ✨augmented speech-to-text model✨ can capture voice-native signals - connecting what someone said with how they sounded - and pass that information along to the AI agent. That means a cascade can preserve the context needed to generate natural responses, without requiring an end-to-end speech-to-speech model. 💡 Natural voice AI doesn't require throwing away the cascade. It requires making sure the cascade doesn't throw away the voice. And done well, that can deliver natural AI conversations in a more cost-effective architecture. 💰💪

  • Bans feel decisive. They're not always effective 👀 Schell Games proved it with Among Us 3D: after redesigning enforcement around ToxMod's real-time voice detection, they cut player disruption by 97.5% - and got better results than banning ever did. A single warning stopped 66.3% of offenders from reoffending. A 1-day ban only stopped 61.4%. Sometimes the smarter fix isn't the harshest one. Full case study in comments 👇

    • No alternative text description for this image
  • $15.6B. That's what account takeover fraud cost US adults in 2024 alone - up 23% year over year 📈 Most fraud tools never heard it happen. 👀 They monitor transactions. They flag metadata. They analyze the transcript after the call ends. All useful — none of it catches fraud while it's happening, because the tells aren't in the words. They're in the voice. Here's what a transcript misses on every fraud call: 1️⃣ Scripted or rehearsed speech patterns 2️⃣ Emotional incongruence: calm under pressure, urgency without stress 3️⃣ Hesitation, contradictions, and recall gaps 4️⃣ Voice mismatch against the account holder 5️⃣ Synthetic or AI-generated voice indicators Modulate's Velma catches all five live - flagging deceptive intent to your agents before the call completes, not after the damage is done. And on Hugging Face's independent leaderboard, Velma's deepfake detector generates fewer than half the errors of the next-best model. Stop reading the transcript. Start listening to the call 💡 See it in action: https://hubs.ly/Q04vYPnT0

  • AI is moving fast. So are the people finding new ways to misuse it. 🤷♀️ That creates a race between attack and defense - and the uncomfortable question is: what happens when attackers start moving faster than our defenses? 👀 For Carter Huffman, that question is at the heart of why we’re building Modulate. As synthetic voices become harder to distinguish from real ones, detection and defense need to evolve just as quickly. The goal isn’t to slow down innovation. It’s to make sure we can keep up with it. Watch Carter share what keeps him up at night, and why it’s driving the work we’re doing at Modulate. 👇

  • 🔉 We’re hiring a Backend Engineer at Modulate. Help build and scale the Developer and Enterprise APIs powering our voice AI models. Ship code that puts ML audio models into production - not just papers. If you’re strong in Python, enjoy building production APIs, and want to work on greenfield systems with real-world impact, we’d love to hear from you! 🙌 Apply here: https://hubs.ly/Q04vH6GZ0

    • No alternative text description for this image
  • Accuracy, speed, and price - the clear winner "without comparison." 🏅 That's how one G2 reviewer described Velma after testing it against every provider on the market. ⭐ 5/5: "Outstanding Accuracy, Speed, and Value" Real developers are shipping faster with Modulate's speech-to-text. Simple API, clear docs, integration in minutes. Test it yourself. Grab 1000 free API credits from the comments 👇

    • No alternative text description for this image
  • 6 reasons Modulate Transcribe is different from every other transcription API on the market: 1️⃣ It's #1 on the Hugging Face Open ASR Leaderboard, out of 88 models evaluated on word error rate across real-world audio. 2️⃣ It's built on Modulate's Ensemble Listening Model (ELM) architecture - orchestrating multiple specialized transcription models together, instead of relying on one, for better accuracy, latency, and cost than any single model alone. 3️⃣ It detects 20+ emotions straight from voice: frustrated, anxious, confident, relieved, and more - so you know how something was said, not just what. 4️⃣ It identifies 12+ accents, from American and British to Indian, Latin American, and Middle Eastern, so accuracy holds up across diverse, global audiences. 5️⃣ No other transcription API surfaces emotion and accent natively from audio - most bolt it on after the fact, if at all. 6️⃣ It posts the #1 score for Mandarin (zh-CN) transcription on Sierra's μ-Bench.

    • No alternative text description for this image
  • We built the detection. Innersloth and Schell Games built a better system with it. 97.5% less player disruption. 🎧 Among Us 3D redesigned enforcement: swapping bans for a tiered escalation system powered by ToxMod's real-time detection. The results: 👉 66.3% of offenders stopped after just ONE warning (better than a 1-day ban) 👉 Ban appeals dropped from 1-in-364 actions to zero 👉 Enforcement disruption fell 97.5% - with zero tradeoff on safety Proof that smarter enforcement beats harsher enforcement. Full case study in the caption👇

Similar pages

Browse jobs