1. X
  2. Albert Gu
Log inSign up
Albert Gu
Cartesia
567 posts
user avatar
Albert Gu
Cartesia
@_albertgu
assistant prof @mldcmu. chief scientist @cartesia_ai. leading the ssm revolution.
Joined December 2018
78
Following
21.3K
Followers
RepliesRepliesMediaMedia
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jul 28
    Evaluations are difficult and vague for all generative models, and benchmarks only capture a small slice. Our blog post dives into the nuances for TTS
    user avatar
    Cartesia
    @cartesia
    Jul 28
    "Is this TTS model good?" gets harder to answer as models improve. "Good" is at least five axes: correctness, naturalness, contextual correctness, robustness, and most evals only capture the first. We wrote up the failure modes that make TTS eval hard: cartesia.ai/blog/is-this-t…
    10K
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jul 8
    Cartesia is hosting an ICML party tomorrow (Thursday) night! it'll go late, but come early or i may have to bounce you again
    Image
    Cartesia's ICML After Party · Luma
    From luma.com
    34K
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jun 26
    Transformers are better at copying, while RNNs are better at modeling "meaning-bearing words—the nouns, verbs, & adjectives that say what a sentence is about"
    user avatar
    Ai2
    @allen_ai
    Jun 25
    Hybrid (transformer–RNN) models are fast becoming a serious alternative to the transformer, but a big question remains: how do they process tokens differently & how does this impact performance? We compared our transformer (Olmo 3) & hybrid (Olmo Hybrid) models to find out. 🧵
    Image
    59K
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jun 19
    Rather than interleaving layers naively, a more fine-grained approach to hybrid models is to allow hybridization across the sequence models within a single layer. The fact that softmax attention and linear attention use similar underlying projection parameters allows switching
    user avatar
    Kevin Li
    @kevinyli_
    Jun 18
    Excited to share last summer's work at Google Research! Most hybrid models today are static: each token sees the same interleaved pattern of your favorite linear model and attention. Oryx instead varies the model used across the sequence through shared representations. 1/
    Image
    32K
  • user avatar
    Albert Gu
    Cartesia
    @_albertgu
    Jun 18
    Congrats to Henry and Naomi - they’ve been so on top of the space and super helpful as collaborators too!
    user avatar
    Henry Yin✈️ICML
    @HenryYin_
    Jun 16
    Most AI investing happens downstream of the frontier: a capability emerges, a category gets named, and capital rushes in. But by the time a category earns a clean box on a market map, the best builders have usually been living in the messy version for months. Agents. Reasoning.
    Image
    9.4K
  • See @_albertgu's full profile

    Sign up
    Log in

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement