Skip to main content

Voxbento: Real-Time Interpretation for Live Events

Voxbento lets you run professional simultaneous interpretation for any live event, entirely in the browser. Interpreters monitor the floor session through a built-in Jitsi console and broadcast translated audio directly to attendees via WebRTC — with sub-second latency and seamless handoff between interpreter shifts.

Quick Start

Set up your first event, booth, and invite interpreters in minutes.

How It Works

Understand the interpretation flow from floor audio to listener playback.

Admin Guide

Create events, manage rooms, assign roles, and generate invite tokens.

API Reference

Integrate Voxbento into your event platform using the REST and WebSocket APIs.

How Voxbento Works

Create an event and booths

Log in to the admin panel and set up your event with rooms and language booths. Each booth represents one interpretation channel (e.g., English, French, Spanish).

Invite interpreters and coordinators

Generate invite tokens for each booth and share the links. Participants open their unique link — no account needed — and join the booth directly in their browser.

Go live

The active interpreter clicks Go Live. Audio streams instantly via WebRTC/WHIP to MediaMTX and is served to listeners with sub-second latency.

Listeners tune in

Attendees visit the listener page for your event, pick their language, and hear the interpretation in real time — with optional live captions powered by AI transcription.

Key Features

Sub-second Latency

WebRTC/WHEP delivery keeps audio delay under one second for every listener.

Seamless Handoff

Coordinators switch active interpreters without dropping listeners or interrupting playback.

Live Captions & AI Translation

Automatic transcription and real-time translation powered by LLMs (OpenAI, Anthropic, Groq, Deepgram, Gemini).

Floor Audio Support

Capture the main stage audio seamlessly with the built-in Floor Bot to provide captions and translations for the original speaker.

No Install Required

Everything runs in the browser. Interpreters and listeners need no plugins or apps.

Multi-language

Run unlimited language booths simultaneously for a single event.

Role-based Access

Fine-grained roles — coordinator, interpreter, listener — scoped per event or per booth.