1. X
  2. Xinyu Yang
Log inSign up
Xinyu Yang
1,133 posts
user avatar

Xinyu Yang

@Xinyu2ML
Building open frontier intelligence. Opinions are my own. They/Them
Mountain View, CA
xinyuyang.me
Joined December 2022
1,567
Following
18K
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Xinyu Yang
    @Xinyu2ML
    Jul 17
    Why can Kimi ship K3? Let me tell my story. Earlier this year, I left academia for industry. I talked to a lot of companies along the way. Here's what I saw: 1⃣Arrogance. They believe the AI war is over, and they won. No hunger for the future, and no hunger for talent.
    Image
    Made with AI
  • user avatar
    Xinyu Yang
    @Xinyu2ML
    Aug 14
    One more reason why you should leave US academia today lol
    user avatar
    Hao Kang
    @GT_HaoKang
    Aug 14
    I wonder if they know how many works and project are conducted by these international interns. For me, GEAR(KV compression, Neurips2024) and TurboAttention(attention acceleration, Mlsys2025) are used in @MSFTResearch coding model service. ThunderAgent(Agentic Infra ICML2026)
    Image
  • user avatar
    Xinyu Yang
    @Xinyu2ML
    Aug 12
    Congrats
    user avatar
    SpaceXAI
    @SpaceXAI
    Aug 12
    Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.
    Image
  • user avatar
    Xinyu Yang
    @Xinyu2ML
    Aug 11
    For sure, mhc/attnres plus looped transformers should be interesting
    user avatar
    Skye
    ZeroEntropy (YC W25)
    @skye7821
    Aug 11
    Replying to @Xinyu2ML
    I feel it may be simpler to use something like hyperconnections or muddformer that expands the residual pathways and improves representation in later layers. I believe there was a paper blending mHC and Looped Transformers recently…
  • user avatar
    Xinyu Yang
    @Xinyu2ML
    Aug 11
    Looped Transformer has been widely discussed this year. However, a key dilemma in prior work remains: if you maintain training parallelism, early layers still cannot access information from later layers. Otherwise, you have to suffer from train-test mismatch. Full-bandwidth
    user avatar
    Xidulu
    @xidulu
    Aug 11
    1/ Sharing a new, interesting project I did during my internship at Microsoft AI Frontiers w/ @JohnCLangford TL;DR: At decoding time, we feed **previous hidden state** into the input together with token embedding, and it boosts performance for free. arxiv.org/abs/2608.08888

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement