Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
- Deepgram has been shipping realtime speech models that deliver both very high accuracy and very low latency since before we had LLMs we could use to build realtime voice agents! Scott played me clips of their new Flux TTS model while they were training it, and I've talked a lotCongrats to @DeepgramAI on the GA launch of Flux TTS, now live in Pipecat 🎉 @JonPTaylor looks at how Flux TTS delivers a more consistent voice experience. Flux reads the whole conversation, not just the next line — adaptive tone, consistent pronunciation, clean interruption
- Developing a robot by chatting with a robot x.com/i/broadcasts/1…
- Jon has been exploring UI for AI agents: voice agents, task agents, systems of multiple agents. Here's an experiment in driving a "the agent is talking" visual component by doing a little bit of formant analysis on the text-to-speech stream.Voice agent UI feels stuck in the wiggly-waveform / glowing-orb phase. High-end avatars are exciting, but simple client-rendered characters need richer data than just amplitude to convey speech. An RMS-driven mouth flap gets you pretty far, but believable mouths need phonetic






