Log inSign up
kwindla
Daily
6,599 posts
kwindla profile banner
@kwindla

kwindla

Daily
@kwindla
Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
San Francisco, CA
machine-theory.com
Joined September 2008
3,924
Following
15.7K
Followers
1
Subscription
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @kwindla
    kwindla
    Daily
    @kwindla
    Aug 27
    Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a
    Image
    00:00
    115
  • @kwindla
    kwindla
    Daily
    @kwindla
    Sep 2
    Here's a complete voice agent tutorial focusing on a super low latency configuration: Pipecat PhoneLLM Alpha 1 on @modal, Deepgram transcription, Cartesia voice. You can grab the code and jus run /setup in Claude or Codex. But Merve's video also dives into the details of
    Image
    00:00
    10
  • @kwindla
    kwindla
    Daily
    @kwindla
    Sep 2
    If you're interested in voice UI, or music, or both, it's worth watching all of this video that Jon just dropped. Voice-controlled, generative music. The interface will be familiar to anyone who has used a digital audio workstation. But also ... deep integration of voice makes
    @JonPTaylor
    Jon Taylor
    Daily
    @JonPTaylor
    Aug 31
    Over a year ago, I posted a video where Gemini and I collaborated on an Ableton Live Session. Music AI is fun! Creating voice-native music software is something I'm really passionate about, so, here is Jamcat. Jamcat's concept is a voice controlled, generative music performance
    Image
    00:00
    1
  • @kwindla
    kwindla
    Daily
    @kwindla
    Sep 1
    I joined the Pipecat TV crew to talk about PhoneLLM: a super low-latency LLM trained on voice agent scenarios. We talked about building agents with a small-ish LLM like this, compared to using a big model. PhoneLLM is based on Nemotron 3 Nano, so it's a 30B-A3B MoE. Prompting a
    Image
    00:00
    8
  • @kwindla
    kwindla
    Daily
    @kwindla
    Aug 31
    Everybody building voice agents obsesses about latency. Every enterprise I talk to wants to run LLMs on their own infrastructure (for data management, regulatory/compliance, and cost reasons). PhoneLLM is a new open weights model trained on voice agent tasks: sub-100ms
    Image
    00:00
    9
Advertisement
Advertisement