1. X
  2. kwindla
Log inSign up
kwindla
Daily
6,518 posts
kwindla profile banner
user avatar

kwindla

Daily
@kwindla
Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
San Francisco, CA
machine-theory.com
Joined September 2008
3,918
Following
14.8K
Followers
1
Subscription
RepliesRepliesArticlesArticlesMediaMedia
  • user avatar
    kwindla
    Daily
    @kwindla
    3h
    I'm a big fan of the Nemotron project at NVIDIA. We've been working with their ASR models, the Nemotron 3 Nano/Super/Ultra models, experimental speech models, and now the Nemotron 3.5 Lightning fast agentic model. The new Nemotron 3.5 Lightning is similar to Nemotron 3 Nano,
    Image
  • user avatar
    kwindla
    Daily
    @kwindla
    6h
    Deepgram has been shipping realtime speech models that deliver both very high accuracy and very low latency since before we had LLMs we could use to build realtime voice agents! Scott played me clips of their new Flux TTS model while they were training it, and I've talked a lot
    user avatar
    Pipecat AI
    Daily
    @pipecat_ai
    12h
    Congrats to @DeepgramAI on the GA launch of Flux TTS, now live in Pipecat 🎉 @JonPTaylor looks at how Flux TTS delivers a more consistent voice experience. Flux reads the whole conversation, not just the next line — adaptive tone, consistent pronunciation, clean interruption
    Image
    00:00
  • user avatar
    kwindla
    Daily
    @kwindla
    8h
    This is living in the future.
    user avatar
    Alexis Gallagher
    @alexisgallagher
    9h
    Developing a robot by chatting with a robot x.com/i/broadcasts/1…
  • user avatar
    kwindla
    Daily
    @kwindla
    14h
    Jon has been exploring UI for AI agents: voice agents, task agents, systems of multiple agents. Here's an experiment in driving a "the agent is talking" visual component by doing a little bit of formant analysis on the text-to-speech stream.
    user avatar
    Jon Taylor
    Daily
    @JonPTaylor
    Jul 16
    Voice agent UI feels stuck in the wiggly-waveform / glowing-orb phase. High-end avatars are exciting, but simple client-rendered characters need richer data than just amplitude to convey speech. An RMS-driven mouth flap gets you pretty far, but believable mouths need phonetic
    Image
    00:00
  • user avatar
    kwindla
    Daily
    @kwindla
    Aug 11
    We've spent a lot of time this year working on subagents support in Pipecat. Why would you want to use subagents? Well, in a voice agent, you can never block the main conversation loop. But your agent will often need to do various kinds of complex work that takes time to
    Image
    00:00

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Image
REPLAY
user avatar
Alexis Gallagher
@alexisgallagher
Developing a robot by chatting with a robot
Advertisement
Advertisement