1. X
  2. Diogo Almeida
Log inSign up
Diogo Almeida
128 posts
user avatar
Diogo Almeida
@CompleteSkeptic
Sane + 🌶️ takes in an insane AI world... AI capabilities researcher: co-created RLHF/ChatGPT @ @openai now trying to right the wrong 🤭 (ceo @typesafeai)
Joined October 2014
101
Following
1,307
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • Pinned
    user avatar
    Diogo Almeida
    @CompleteSkeptic
    Jul 24
    Article cover image
    Article
    Is it even possible for the Chinese Labs to distill US models?
    TL;DR: yes, but how it works is unintuitive! Everyone is arguing whether Kimi K3 distilled Anthropic. This whole conversation involves a horrible abuse of the term “distillation” which has come to...
  • user avatar
    Diogo Almeida
    @CompleteSkeptic
    Aug 5
    this has always been the case! reading papers seems to be an early career thing - eventually you get the perspective that (1) most new things don't work and that (2) focus beats chasing new shinies
    user avatar
    Keller Jordan
    @kellerjordan0
    Aug 4
    PSA: Most biglab people now read almost zero papers and understand ICLR/ICML/NeurIPS to be mainly full of overclaims & fraud. (but there are a few diamonds in the rough of course)
  • user avatar
    Diogo Almeida
    @CompleteSkeptic
    Aug 4
    This was a problem in the 2007 Mathematical Contest in Modeling (MCM) (think of it as doing real world ML before it was cool) What makes it so fun and hard is that there are so many additional constraints: - what if a family is boarding together - what if how long a person takes
    user avatar
    Rod
    @rod_mallo
    Aug 4
    boarding was solved mathematically in 2008. no airline has used it once. why?
    Image
    00:00
  • user avatar
    Diogo Almeida
    @CompleteSkeptic
    Aug 1
    gonna sound like an old-timer but I don't think people realize how bad google was at instruction following at the time - their best technique (FLAN) actually hurt performance
    user avatar
    Cheng Lou
    @_chenglou
    Aug 1
    I sometime think about that Jeff Dean interview where he said they had an internal bot before ChatGPT but didn't think it was better than just googling
  • user avatar
    Diogo Almeida
    @CompleteSkeptic
    Jul 31
    I cannot possibly disagree more! This completely ignores the 2nd most important thing to ML impact: data. TL;DR of the RLHF paper is that data >> scale (even SFT closes most of the gap)
    Image
    user avatar
    Horace He
    Thinking Machines
    @cHHillee
    Jul 29
    Presenting my grand unified theory of ML researcher impact: Your impact is directly proportional to how much pain you cause to infra. Fundamentally, you can only inflict pain upon infra if your approach actually works. And the better your approach works the more pain infra is

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement