1. X
  2. Albert Gu
Log inSign up
Albert Gu
Cartesia
575 posts
@_albertgu

Albert Gu

Cartesia
@_albertgu
assistant prof @mldcmu. chief scientist @cartesia. leading the ssm revolution.
Joined December 2018
78
Following
21.4K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @_albertgu
    Albert Gu
    Cartesia
    @_albertgu
    Aug 27
    sonic-3.6 is out, with a large improvement over (the already #1) sonic-3.5 in just a few months! the research team's focus on fundamentals is accelerating progress at the frontier of architectures and audio
    @cartesia
    Cartesia
    @cartesia
    Aug 27
    Sonic-3.6 is now generally available. In January we made a bet: stop tuning the existing paradigm, rebuild from the architecture up. How we topped our own best model in two months → cartesia.ai/blog/sonic-3.6…
    Image
    00:00
    2
  • @_albertgu
    Albert Gu
    Cartesia
    @_albertgu
    Aug 17
    the team continues cooking 👩‍🍳 this is now an unprecedented gap on both the super competitive Provider Voices leaderboard as well as the newer Controlled Voices leaderboard (a stronger benchmark that can't be benchmaxxed, requiring truly better algorithms) Cartesia's TTS model is
    Image
    Image
    Image
    @ArtificialAnlys
    Artificial Analysis
    @ArtificialAnlys
    Aug 17
    Cartesia's Sonic 3.6 takes the #1 spot on both the Provider Voice and Controlled Voice Artificial Analysis Speech Arena leaderboards, surpassing Speechify AI's Simba 3.2 and Alibaba's Qwen-Audio-3.0-TTS-Plus, with Sonic 3.5 holding #2 on Controlled Voice Sonic 3.6 is the latest
    5
  • @_albertgu
    Albert Gu
    Cartesia
    @_albertgu
    Jul 28
    Evaluations are difficult and vague for all generative models, and benchmarks only capture a small slice. Our blog post dives into the nuances for TTS
    @cartesia
    Cartesia
    @cartesia
    Jul 28
    "Is this TTS model good?" gets harder to answer as models improve. "Good" is at least five axes: correctness, naturalness, contextual correctness, robustness, and most evals only capture the first. We wrote up the failure modes that make TTS eval hard: cartesia.ai/blog/is-this-t…
    3
  • @_albertgu
    Albert Gu
    Cartesia
    @_albertgu
    Jul 8
    Cartesia is hosting an ICML party tomorrow (Thursday) night! it'll go late, but come early or i may have to bounce you again
    Image
    Cartesia's ICML After Party · Luma
    From luma.com
    10
  • @_albertgu
    Albert Gu
    Cartesia
    @_albertgu
    Jun 26
    Transformers are better at copying, while RNNs are better at modeling "meaning-bearing words—the nouns, verbs, & adjectives that say what a sentence is about"
    @allen_ai
    Ai2
    @allen_ai
    Jun 25
    Hybrid (transformer–RNN) models are fast becoming a serious alternative to the transformer, but a big question remains: how do they process tokens differently & how does this impact performance? We compared our transformer (Olmo 3) & hybrid (Olmo Hybrid) models to find out. 🧵
    Image
    6
Advertisement
Advertisement