Log inSign up
Nicholas Roberts
626 posts
Nicholas Roberts profile banner
@nick11roberts

Nicholas Roberts

@nick11roberts
Working on foundation models and scaling laws. Postdoctoral research fellow @PrincetonPLI, PhD @WisconsinCS, prev. CMU @mldcmu, UCSD @ucsd_cse, FCC @fresnocity.
New York, NY
nick11roberts.science
Joined April 2012
1,949
Following
1,550
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @nick11roberts
    Nicholas Roberts
    @nick11roberts
    Apr 6
    That new LFM2.5-350M is super overtrained, right? And everyone was shocked about how far they pushed it? As it turns out, we have a brand new scaling law for that! 🧵 [1/n]
    Image
    11
  • @nick11roberts
    Nicholas Roberts
    @nick11roberts
    Sep 4
    I had a wonderful time visiting to give this talk!!! Thank you @datologyai for having me!
    @datologyai
    DatologyAI
    @datologyai
    Sep 4
    A 350M model overtrained to 80,000+ tokens per parameter should be a mistake. Nicholas Roberts' ( @nick11roberts ) scaling law says it's optimal. The Train-to-Test (T²) law folds pass@k into pretraining scaling, so once a model reasons at test time, being compute-optimal flips
    Image
    00:00
  • @nick11roberts
    Nicholas Roberts
    @nick11roberts
    Aug 19
    Great read and congrats on GLM-5.3! Here’s my paper that @jietang cites—optimal tokens/param is task-dependent: memorization wants params, reasoning wants data arxiv.org/abs/2503.10061 He also mentions inference/reasoning cost+scaling. We do this with T² arxiv.org/abs/2604.01411
    @jietang
    jietang
    Z.ai
    @jietang
    Aug 19
    Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have,
    3
  • @nick11roberts
    Nicholas Roberts
    @nick11roberts
    Aug 17
    Super excited to present my work on train-to-test scaling laws today at the Snorkel Reading Group!!! 📈
    @SnorkelAI
    Snorkel AI
    @SnorkelAI
    Aug 17
    Two Reading Groups, one week 👀 Monday, August 17 at 4:00 p.m. PT: @nick11roberts (Incoming Postdoctoral Research Fellow, @Princeton) will present Train-to-Test (T²) Scaling Laws, exploring how inference costs can radically shift optimal pretraining into the overtraining regime.
    Image
  • @nick11roberts
    Nicholas Roberts
    @nick11roberts
    Jul 11
    Asher Trockman.
    @ashertrockman
    Asher Trockman
    @ashertrockman
    Jul 9
    check out this great work by @YixuanEvenXu and @jwkirchenbauer
Advertisement
Advertisement