Voxbento: Real-Time Interpretation for Live Events
Voxbento lets you run professional simultaneous interpretation for any live event, entirely in the browser. Interpreters monitor the floor session through a built-in Jitsi console and broadcast translated audio directly to attendees via WebRTC — with sub-second latency and seamless handoff between interpreter shifts.
Quick Start
Set up your first event, booth, and invite interpreters in minutes.
How It Works
Understand the interpretation flow from floor audio to listener playback.
Admin Guide
Create events, manage rooms, assign roles, and generate invite tokens.
API Reference
Integrate Voxbento into your event platform using the REST and WebSocket APIs.
How Voxbento Works
Create an event and booths
Log in to the admin panel and set up your event with rooms and language booths. Each booth represents one interpretation channel (e.g., English, French, Spanish).
Invite interpreters and coordinators
Generate invite tokens for each booth and share the links. Participants open their unique link — no account needed — and join the booth directly in their browser.
Go live
The active interpreter clicks Go Live. Audio streams instantly via WebRTC/WHIP to MediaMTX and is served to listeners with sub-second latency.
Listeners tune in
Attendees visit the listener page for your event, pick their language, and hear the interpretation in real time — with optional live captions powered by AI transcription.
Key Features
Sub-second Latency
WebRTC/WHEP delivery keeps audio delay under one second for every listener.
Seamless Handoff
Coordinators switch active interpreters without dropping listeners or interrupting playback.
Live Captions & AI Translation
Automatic transcription and real-time translation powered by LLMs (OpenAI, Anthropic, Groq, Deepgram, Gemini).
Floor Audio Support
Capture the main stage audio seamlessly with the built-in Floor Bot to provide captions and translations for the original speaker.
No Install Required
Everything runs in the browser. Interpreters and listeners need no plugins or apps.
Multi-language
Run unlimited language booths simultaneously for a single event.
Role-based Access
Fine-grained roles — coordinator, interpreter, listener — scoped per event or per booth.