1. X
  2. Yi Zhang
Log inSign up
Yi Zhang
60 posts
Image
user avatar
Yi Zhang
@YiZhangZZZ
plant fruits@meta Prev. @xAI, @Apple, @MSFTResearch, PhD @princeton
Menlo Park, CA
yi-zhang.me
Joined June 2022
269
Following
1,395
Followers
RepliesRepliesMediaMedia

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

TermsยทPrivacyยทCookiesยทAccessibilityยทAds Infoยทยฉ 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Jul 8, 2025
    stay tuned for the smartest AI humanity's ever seen๐Ÿš€
    user avatar
    Elon Musk
    X
    @elonmusk
    Jul 7, 2025
    Grok 4 release livestream on Wednesday at 8pm PT @xai
    175K0175K
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Dec 16, 2022
    Very nice visualization by @TheGradient of an example from our paper on 'edge of stability' (arxiv.org/abs/2212.07469). Gradient flow gets trapped in sharp local minima, but the discontinuous dynamics of the 'edge of stability' escape them and settle in flat regions.
    user avatar
    Hossein Mobahi
    @TheGradient
    Dec 16, 2022
    Replying to @SebastienBubeck
    Intriguing work! Not sure if this a coincidence, but plotting the loss in your task (Eq 1.2) reveals a narrow ravine & a wide minimum. The solution that generalizes (i.e. learns threshold neuron) is the flat one. Looks like you cannot realize a threshold neuron at sharper ravine.
    Image
    4K04K
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Apr 5, 2025
    LLMs are sensitive to changes in prompt style โ€” even the best ones, like GPT-4o. In this work with my amazing intern @VectorZhou, we may have found a cure.
    user avatar
    Runlong Zhou
    @VectorZhou
    Apr 4, 2025
    ๐Ÿง  Ever notice how LLMs struggle with familiar knowledge in unfamiliar formats? Our new paper "CASCADE Your Datasets for Cross-Mode Knowledge Retrieval of Language Models" tackles this head-on! ๐Ÿ” Our findings: - Created a qualitative pipeline demonstrating problem we call
    Image
    Image
    1.5K01.5K
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Dec 15, 2023
    Many thanks @BingbinL! for your amazing work๐Ÿซก
    user avatar
    Bingbin Liu
    @BingbinL
    Dec 15, 2023
    Please come to our poster at the MATH-AI workshop! ๐Ÿ˜Š (Room 217-219, 3-4pm)
    1.4K01.4K
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Mar 24, 2023
    Replying to @martin_gorner
    Hi Martin, thanks for your great appreciation of this example. Yuanzhi and I are in charge of the coding section. The specific example's credit goes to Yuanzhi, but since he has yet to claim a Twitter account, please allow me to try to answer the questions (I'll do it inline).
    4340434
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Jul 5, 2022
    The GitHub repo for LEGO is finally online! Let us know how well your favorite transformer does on it
    user avatar
    Sebastien Bubeck
    @SebastienBubeck
    Jul 5, 2022
    After the paper arxiv.org/abs/2206.04301 & the vid youtube.com/watch?v=brmidgโ€ฆ, here comes the Github repo for LEGO, courtesy of our amazing Yi Zhang @YiZhangZZZ! Lots more can be uncovered abt Transformers with LEGO. Play with it & let us know what you find! github.com/yizhangzzz/traโ€ฆ
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Jun 7, 2025
    Replying to @SebastienBubeck
    Congrats Seb!
    4500450
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Jul 7, 2022
    The Scientist Who Developed a New Way to Understand Communication quantamagazine.org/mark-bravermanโ€ฆ via @QuantaMagazine
    quantamagazine.org
    Mark Braverman Wins the IMU Abacus Medal | Quanta Magazine
    Mark Braverman has spent his career translating thorny problems into the language of information complexity.
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Dec 16, 2022
    Replying to @TheGradient and @SebastienBubeck
    That's correct, and nice plot btw! Our interpretation of this example is that the edge of stability leads to a discontinuous trajectory that escapes from the bad local minima (i.e. the narrow ravine) while continuous (or small learning rate) gradient dynamics will get trapped.
    4360436
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Jul 31, 2022
    Replying to @SebastienBubeck
    wow, congrats Seb๏ผ
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Dec 16, 2022
    Replying to @deepcohen @TheGradient and @SebastienBubeck
    I mean, the *test loss* is worse for the gradient flow trajectory. More conceptually, from the same initialization, gradient flow does not learn the threshold neuron while EOS does.
    1990199
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Mar 24, 2023
    Replying to @martin_gorner and @erichorvitz
    Let me know if I was able to answer your questions above. Followups are welcomed ๐Ÿ˜ƒ
    44044
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Mar 24, 2023
    Replying to @martin_gorner
    Yuanzhi's idea for this line of prompt is that given a scalar hyper parameter d_dim, and each gradient, which is a tensor of any order, has an axis whose dimension is exactly d_dim. Find that axis ('by looping') and then reshape the gradient to 2d by collapsing all the other axes
    42042
  • user avatar
    Yi Zhang
    @YiZhangZZZ
    Mar 24, 2023
    Replying to @martin_gorner
    you are absolutely right here, this is technically an incorrect answer from the model. We will highlight it together with the other mistakes in the paper.
    35035
Advertisement
Advertisement