See Meta's launch announcement, architecture notes, adaptive-delay explanation, language coverage, speaker demos, and long-context examples at the primary source.
Muse Voice Transcribe: Learn the Model and Try Speech to Text Online
Muse Voice Transcribe is Meta Superintelligence Labs' real-time audio perception model for streaming speech recognition, speaker diarization, and endpoint detection.
Read what Meta published about its multilingual model, then use our browser workspace to upload audio or video, record on the page, or import a media link.
Review timestamped text, rename speakers, and export the corrected transcript as TXT, DOCX, PDF, SRT, VTT, or JSON.
AI speech to text
Audio intake
What Is Meta Muse Voice Transcribe?
Meta introduced Muse Voice Transcribe on September 1, 2026 as a real-time audio perception model. Meta says it combines streaming ASR, diarization for 20+ speakers, and endpointing; it was trained across 70+ languages, with 25 extensively verified, and supports code-switching plus language, keyword, and context biasing. Those capabilities describe Meta's model. This page also gives you a practical browser workflow for turning your own media into reviewable text.

Open the official developer page when you need the Muse Voice Transcribe API, model access details, or integration documentation for your own application.
Upload audio or video up to 1 GB, record on the page, or paste a hosted media link. Then review timed text, rename speakers, and export the corrected version in six formats.
From Muse Voice Research to a Transcript You Can Use
If you arrived to understand Muse Voice Transcribe, the official sources above explain the model. If you need a finished transcript, the browser workspace takes your recording through timed review, speaker cleanup, and export without asking you to wire an API first.

Speaker Labels and Word-Level Timestamps
Turn on speaker separation and the app splits the conversation into turns, then lets you rename Speaker 1 into a real name across the whole file. Every word carries a timestamp you can click to hear that moment again, which settles a disputed phrase in seconds.

Meetings, Interviews, Podcasts, Lectures, Captions
People run Muse Voice speech to text over meeting recordings for decision logs, interviews for journalism and qualitative research, podcast episodes for show notes, lectures for study notes, video files for YouTube subtitles, and recorded talks for accessibility captions.

An Editor, Not Just a Download Button
The first pass is a draft. Search it for a phrase, fix a product name the model has never heard, correct the punctuation, rename the speakers, then export. Every file you download is built from the version you corrected, never the untouched model output.
Use Muse Voice as a Complete Browser Workflow
Move from source media to a reviewed handoff in one place: what the workspace accepts, how you check the draft, and what you can export afterwards.
Upload Audio or Video Up to 1 GB
Send MP3, WAV, M4A or FLAC for audio and MP4 or MOV for video, up to 1 GB per file. Audio to text and video to text run down the same path here, so you never strip a soundtrack out of a clip first.
Record Without Leaving the Page
Press record in the browser and talk: a stand-up, an interview on speakerphone, a voice memo. The capture goes straight into the queue, skipping the usual save-it, find-it, upload-it detour.
Start From a Hosted Media Link
Already have the recording hosted somewhere with a direct media address? Paste that URL and we fetch it, which beats pulling a large file down just to push the same bytes back up again.
About 100 Languages, Detected or Chosen
Let the app detect the spoken language, or select it yourself when you already know. Accuracy depends on audio quality, accents and background noise, and it varies by language, so clean close-miked speech needs the least work.
Search, Correct, Rename Speakers
Find any phrase in the transcript, fix a misheard term, and change a generic speaker label everywhere it appears. Corrections save against the transcript, so whatever you export next carries them.
Six Exports, Captions Included
Take the finished transcript as TXT or DOCX for writing, PDF for a fixed reading copy, SRT or VTT for subtitles and closed captions, or JSON when something downstream needs the timings intact.
Muse Voice vs Other Online Transcription Tools
Muse Voice brings file upload, browser recording, hosted media import, timed review, renameable speakers, and six export formats into one browser workflow. Compare that complete path with tools built mainly for meetings, human-reviewed transcripts, video editing, or high-volume uploads.
| Tool | Ways to start | Language coverage | Speaker workflow | Review & edit | Export options | Best for |
|---|---|---|---|---|---|---|
| Muse Voice | Upload audio or video, record in the browser, or paste a hosted media URL | About 100 languages; auto-detect or choose | Optional separation; rename a label across the transcript | Search, correct, and replay from word timestamps | TXT, DOCX, PDF, SRT, VTT, and JSON | A direct browser path from recording to reviewed handoff |
| Otter.ai | Record meetings or import audio and video | Six transcription languages | Automatic speaker identification and names | Search, edit, highlight, and create meeting notes | TXT, DOCX, PDF, SRT, and audio | Live meeting notes, summaries, and team follow-up |
| Rev | Upload recorded audio or video | AI and human services across supported languages | Speaker labels; human name reconciliation is available | Online review with human-verified options | TXT, SRT, VTT, and professional caption formats | Human-reviewed transcripts and accessibility captions |
| Descript | Upload or import media, or record in Descript | Transcription and translation across supported languages | Automatic speaker detection and labels | Edit the transcript and media on the same timeline | Text, Word, SRT, VTT, video, and more | Transcript-led audio and video editing |
| TurboScribe | Upload common audio or video files | 98+ languages | Optional speaker recognition | Edit transcripts and translate speech to English | SRT, VTT, and transcript downloads | High-volume files and very long recordings |
| Notta | Upload, record live, or import supported links | 58 transcription languages | Speaker identification on supported workflows | Edit, search, replay, and collaborate | TXT, DOCX, PDF, SRT, XLSX, and sharing | Meeting notes across web, mobile, and extensions |
Feature details last verified:
Feature details were verified against each vendor's official product or help pages on 2026-09-03. Availability can vary by plan and change over time.
Bring audio in three ways
Upload an audio or video file, record directly in the page, or paste a hosted media URL. Start from the source you already have instead of reshaping it for the tool.
Review against the exact moment
Search the transcript and jump back from a timestamp to the matching audio. Names, numbers, and disputed phrases can be checked without scrubbing through the whole recording.
Rename a speaker once
Replace a generic label across every matching turn so the reviewed transcript and every later export keep the same speaker names.
Export one corrected transcript six ways
Create TXT, DOCX, PDF, SRT, VTT, or JSON from the version you reviewed. A correction made once carries into documents, captions, and structured data.
Use the workflow without installing or wiring an API
The browser workspace connects intake, transcription, review, and export. You can test the process on your own recording before deciding whether you need a developer integration.
Muse Voice Transcription Pricing
Choose minutes by the way your work arrives: a recurring allowance for a steady recording schedule or a one-time pack for an occasional backlog. New accounts can test the workflow with 5 free minutes, and the same balance also covers the available AI tools.
Starter
For a predictable stream of short interviews, lessons, or weekly calls, paid annually.
Allowance and workflow
- Annual allowance: 1,440 transcription minutes
- Monthly equivalent: 120 minutes
- One annual charge of $58.80
- Accepts audio or video uploads and recordings up to 1GB
- Timed text with optional speaker labelling
- Edit the transcript before delivery
- Create six outputs from your approved transcript: TXT, SRT, VTT, JSON, PDF, or DOCX
- Available AI tools use credits without requiring a Pro or Max plan
The displayed monthly figure is an equivalent; checkout bills the full year.
Pro
Best valueA lower subscription rate per minute for teams or creators with about ten hours of media each month.
Allowance and workflow
- Annual allowance: 7,200 transcription minutes
- Monthly equivalent: 600 minutes
- One annual charge of $178.80
- The complete Starter workflow is included
- Spend credits on the available AI tools without another plan gate
The price shown per month is an equivalent; the annual total is charged together.
Max
Built for large archives and recurring batches that add up to roughly fifty hours each month.
Allowance and workflow
- Annual allowance: 36,000 transcription minutes
- Monthly equivalent: 3,000 minutes
- One annual charge of $298.80
- Includes the complete Pro transcription toolset
- Lowest subscription cost per minute for high-volume use
- Built for reviewing a higher volume of recorded minutes
- Account storage for private audio
- Use credits for the available AI tools without a separate tier
Checkout charges the annual amount; the monthly number is provided for comparison.
Muse Voice Transcribe: Questions Before You Try It
Official model facts, what the browser workspace can do today, and the limits to check before you rely on a transcript.
Click questions to expand detailed answers
Muse Voice Transcribe is Meta Superintelligence Labs' real-time audio perception model. Meta says it combines streaming automatic speech recognition, speaker diarization, and endpointing, with multilingual and code-switching support plus language, keyword, and context biasing.
Try the Transcription Workflow on Your Own Recording
Upload a file, record on the page, or paste a link. Then review the timed draft, correct a line, rename a speaker, and export the same approved text as a document, caption file, or structured data.
