1. X
  2. Yangjun Ruan
Log inSign up
Yangjun Ruan
253 posts
user avatar
Yangjun Ruan
@YangjunR
Compressing
Palo Alto, CA
cs.toronto.edu/~yjruan/
Joined February 2021
807
Following
1,437
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Yangjun Ruan
    @YangjunR
    Mar 26, 2025
    New paper on synthetic pretraining! We show LMs can synthesize their own thoughts for more data-efficient pretraining, bootstrapping their capabilities on limited, task-agnostic data. We call this new paradigm “reasoning to learn”. arxiv.org/abs/2503.18866 Here’s how it works🧵
    Image
    52K
  • user avatar
    Yangjun Ruan
    @YangjunR
    Jan 23
    I always think TTT as the best scientific setup for studying data efficiency in the limit - and here we have some signs of life that there are very data-efficiency learning paradigms
    This post is from a protected account.
    4.6K
  • user avatar
    Yangjun Ruan
    @YangjunR
    Jan 12
    We've seen pretraining as such a powerful learning paradigm by compressing information in the context into weights - now we should start doing that at test time, too.
    user avatar
    Karan Dalal
    @karansdalal
    Jan 12
    LLM memory is considered one of the hardest problems in AI. All we have today are endless hacks and workarounds. But the root solution has always been right in front of us. Next-token prediction is already an effective compressor. We don’t need a radical new architecture. The
    Image
    1.5K
  • user avatar
    Yangjun Ruan
    @YangjunR
    Dec 2, 2025
    I’ll be attending #NeurIPS starting Wednesday as part of @thinkymachines! Feel free to DM me if you’d like to catch up, chat about research, or learn more about Thinky (we have openings!)🤝 job-boards.greenhouse.io/thinkingmachin…
    17K
  • user avatar
    Yangjun Ruan
    @YangjunR
    Nov 21, 2025
    Observational scaling laws hold!
    user avatar
    Epoch AI
    @EpochAIResearch
    Nov 20, 2025
    Benchmarking data is dominated by a single “General Capability” dimension. Is this due to good generalization across tasks, or to developers pushing on all benchmarks at once? 🧵 with some analysis, including the discovery of a “Claudiness” dimension.
    Image
    804

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement