Log inSign up
David Alvarez Melis
855 posts
David Alvarez Melis profile banner
@elmelis

David Alvarez Melis

@elmelis
Asst. Prof. @hseas @KempnerInst || Researcher @MSRNE || ML + NLP || Previously: @MIT_CSAIL NYU @IBMResearch @ITAM_mx
Cambridge, MA
dmelis.github.io
Joined August 2010
2,327
Following
2,047
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @elmelis
    David Alvarez Melis
    @elmelis
    Aug 11
    Very nice project led by @sunnytqin and @kimiahmdh! Classical scaling laws implicitly assumed data was free, but nothing in life is free😆. So the right question is what is the *exchange rate* between fresh and derived tokens, and our results show that rate is far from constant.
    @sunnytqin
    Sunny Qin
    @sunnytqin
    Aug 10
    (1/N) 🧵 Chinchilla assumes you'll never run out of fresh data. That era is ending! Compute keeps growing exponentially, but high-quality tokens don't. So what's the exchange rate between extra compute and fresh, high-quality data? We propose Compute-Data (CD) scaling laws,
    Image
    1
  • @elmelis
    David Alvarez Melis
    @elmelis
    Jul 16
    Very excited about this work, led brilliantly by @kimiahmdh Does your data mixture give you synergy vibes? Like math-and-code-kind-of-vibes? Here's one (technically, two) ways to turn those vibes into concrete estimates, which can then be used to optimize mixture design.
    @kimiahmdh
    Kimia Hamidieh
    @kimiahmdh
    Jul 16
    Can we tell whether data domains cooperate or compete during pretraining? Adding code to the mix makes models better at math, while some other combinations hurt each other. We call this data synergy. Turns out you can incorporate data synergy into scaling laws and estimate it 🧵
    Image
  • @elmelis
    David Alvarez Melis
    @elmelis
    Jun 5
    This was a fun one! And a real treat to be a (small) part of it. Partly because the paper formalizes a bunch of things about scale/data that felt plausible but fuzzy to me before, and partly because watching Ekdeep in action is a treat of its own. He's one of a kind.
    @ChrisGPotts
    Christopher Potts
    @ChrisGPotts
    Jun 1
    We take for granted that larger models are better than smaller ones, but why is this so? Our new paper, led by Jing Huang and @EkdeepL, traces this to a data-induced competition for resources (neurons), using formal analysis, idealized tasks, and real pretraining.
    Title card for a research paper. The title reads "Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention." Authors listed: Jing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra, David Alvarez-Melis, Andrew Lampinen, Christopher Potts, and Ekdeep Singh Lubana. A Goodfire logo appears below the names. Author affiliations: Stanford University, Kempner Institute at Harvard University, MIT, and Anthropic.
    1
  • @elmelis
    David Alvarez Melis
    @elmelis
    Apr 23
    Our Data-Centric ML group is at ICLR 🇧🇷this week. I couldn't make it this year 😰, but @SaraKangaslahti, @JonathanGeuter, @rach_it_ are there. Find them, say hi. Quick rundown 👇
    1
  • @elmelis
    David Alvarez Melis
    @elmelis
    Feb 18
    The terrific trio of @rach_it_ @clara_mohri @sunnytqin just dropped a blog on “RL excursions” during LLM pretraining. Context: Lots of recent work tries to bring RL earlier intro pretraining with fancy new methods/training paradigms…
    @rach_it_
    Rachit Bansal
    @rach_it_
    Feb 18
    In standard LLM training, RL comes last. In our new work, we question this paradigm. So, when does an LLM become capable of learning via RL? Short answer: Much earlier than you expect! Blogpost: rl-excursions.github.io w/ @clara_mohri @sunnytqin @elmelis @ShamKakade6 🧵(1/n)
    Image
    GIF
    1
Advertisement
Advertisement