Log inSign up
Rosinality
33.3K posts
Rosinality profile banner
@rosinality

Rosinality

@rosinality
ML Engineer @poolsideai
London, United Kingdom
github.com/rosinality
Joined October 2008
1,000
Following
8,007
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @rosinality
    Rosinality
    @rosinality
    Feb 5
    I post the papers I find interesting. There are so many papers published these days, and I frequently miss great papers. I appreciate paper recommendations via DM, but I tend to only post papers I discover on my own to keep my list personally curated.
    6
  • @rosinality
    Rosinality
    @rosinality
    Sep 5
    I can't understand why people keep trying to say some company won the race after each model release. What is important for the model company is whether they have a roadmap and good direction and whether they are able to achieve it, not the model at each specific time point which
    1
  • @rosinality
    Rosinality
    @rosinality
    Sep 2
    Maybe full-bandwidth transformer (not looped transformer) style architecture could allow more obscure cot? Though I think it will still be anchored around discrete tokens.
    4
  • @rosinality
    Rosinality
    @rosinality
    Sep 2
    Great results when everyone talks about looped transformers. Looped transformers are now more compute-efficient compared to non-looped ones.
    @wangsw5653
    Shaowen Wang
    @wangsw5653
    Sep 2
    Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: arxiv.org/abs/2609.01343
    Image
    2
  • @rosinality
    Rosinality
    @rosinality
    Sep 1
    arxiv.org/abs/2608.30627 Inserting reasoning tokens into the pretraining data. This has been tried multiple times, but how scalable is it?
    Image
    Image
    4
Advertisement
Advertisement