1. X
  2. Albert Gu
Log inSign up
Albert Gu
Cartesia
571 posts
user avatar

Albert Gu

Cartesia
@_albertgu
assistant prof @mldcmu. chief scientist @cartesia. leading the ssm revolution.
Joined December 2018
78
Following
21.4K
Followers
RepliesRepliesMediaMedia
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Aug 17
    the team continues cooking 👩‍🍳 this is now an unprecedented gap on both the super competitive Provider Voices leaderboard as well as the newer Controlled Voices leaderboard (a stronger benchmark that can't be benchmaxxed, requiring truly better algorithms) Cartesia's TTS model is
    Image
    Image
    Image
    user avatar
    Artificial Analysis
    @ArtificialAnlys
    Aug 17
    Cartesia's Sonic 3.6 takes the #1 spot on both the Provider Voice and Controlled Voice Artificial Analysis Speech Arena leaderboards, surpassing Speechify AI's Simba 3.2 and Alibaba's Qwen-Audio-3.0-TTS-Plus, with Sonic 3.5 holding #2 on Controlled Voice Sonic 3.6 is the latest
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jul 28
    Evaluations are difficult and vague for all generative models, and benchmarks only capture a small slice. Our blog post dives into the nuances for TTS
    user avatar
    Cartesia
    @cartesia
    Jul 28
    "Is this TTS model good?" gets harder to answer as models improve. "Good" is at least five axes: correctness, naturalness, contextual correctness, robustness, and most evals only capture the first. We wrote up the failure modes that make TTS eval hard: cartesia.ai/blog/is-this-t…
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jul 8
    Cartesia is hosting an ICML party tomorrow (Thursday) night! it'll go late, but come early or i may have to bounce you again
    Image
    Cartesia's ICML After Party · Luma
    From luma.com
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jun 26
    Transformers are better at copying, while RNNs are better at modeling "meaning-bearing words—the nouns, verbs, & adjectives that say what a sentence is about"
    user avatar
    Ai2
    @allen_ai
    Jun 25
    Hybrid (transformer–RNN) models are fast becoming a serious alternative to the transformer, but a big question remains: how do they process tokens differently & how does this impact performance? We compared our transformer (Olmo 3) & hybrid (Olmo Hybrid) models to find out. 🧵
    Image
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jun 19
    Rather than interleaving layers naively, a more fine-grained approach to hybrid models is to allow hybridization across the sequence models within a single layer. The fact that softmax attention and linear attention use similar underlying projection parameters allows switching
    user avatar
    Kevin Li
    @kevinyli_
    Jun 18
    Excited to share last summer's work at Google Research! Most hybrid models today are static: each token sees the same interleaved pattern of your favorite linear model and attention. Oryx instead varies the model used across the sequence through shared representations. 1/
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement