Log inSign up
Chao Du
AMI Labs
15 posts
@duchao0726

Chao Du

AMI Labs
@duchao0726
Building world models @amilabs
Singapore
duchao0726.github.io
Joined January 2011
171
Following
594
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @duchao0726
    Chao Du
    AMI Labs
    @duchao0726
    Mar 10
    Understanding the real world is key to building advanced AI systems. Excited to join @amilabs at launch with a brilliant team to make it happen!
    @amilabs
    AMI Labs
    @amilabs
    Mar 10
    Advanced Machine Intelligence (AMI) is building a new breed of AI systems that understand the world, have persistent memory, can reason and plan, and are controllable and safe. We’ve raised a $1.03B (~€890M) round from global investors who believe in our vision of universally
    Photographer: Yann LeCun

IC1340 / NGC 6992 / Eastern Veil nebula
20210628-ic1340-rasa-2600mc-lext
Scope: Celestron RASA 11"
Camera: ZWO ASI2600MC
Filter: Radian Triad quad narrow band.
Subs: 65 @300 seconds.
    8
  • @duchao0726
    Chao Du
    AMI Labs
    @duchao0726
    Nov 28, 2025
    Interesting analysis! On the log-sum vs. sum-log: the DLM training objective, from a VI perspective, essentially forces posterior collapse (over ordering). This shifts burden from inference to generation. It enables things like infilling but also increases modeling difficulty.
    @ducx_du
    Cunxiao Du
    @ducx_du
    Nov 25, 2025
    Diffusion LLMs (DLLM) can do “any-order” generation, in principle, more flexible than left-to-right (L2R) LLM. Our main finding is uncomfortable: ➡️ In real language, this flexibility backfires: DLLMs become worse probabilistic models than the L2R / R2L AR LMs. This
  • @duchao0726
    Chao Du
    AMI Labs
    @duchao0726
    Oct 31, 2025
    Our new work on RL training–inference mismatch shows that simply reverting to FP16 makes RL much more stable across algos, models, and engines. Feels like we’re reaching a stage where we can rethink RL for LLMs again with a clean off-policy PG formulation.
    @rosinality
    Rosinality
    @rosinality
    Oct 31, 2025
    FP16 can have a smaller training-inference gap compared to BFloat16, thus fits better for RL. Even the difference between RL algorithms vanishes once FP16 is adopted. Surprising!
    Image
    1
  • @duchao0726
    Chao Du
    AMI Labs
    @duchao0726
    May 28, 2025
    An elegant way to learn from supervised data, via reinforcement learning, to elicit reasoning behaviors in LLMs.
    @zzlccc
    Zichen Liu
    @zzlccc
    May 28, 2025
    Reinforcing General Reasoning without Verifiers 🈚️ R1-Zero-like RL thrives in domains with verifiable rewards (code, math). But real-world reasoning (chem, bio, econ…) lacks easy rule-based verifiers — and model-based verifiers add complexity. Introducing *VeriFree*: ⚡ Skip
    Image
  • @duchao0726
    Chao Du
    AMI Labs
    @duchao0726
    Feb 7, 2025
    Sharing interesting findings in R1-Zero-like training.
    Image
    @zzlccc
    Zichen Liu
    @zzlccc
    Feb 6, 2025
    🚨There May Not be Aha Moment in R1-Zero-like Training: oatllm.notion.site/oat-zero A common belief about the recent R1-Zero-like training is that self-reflections *emerge* as a result of RL training. We carefully investigated and showed the opposite. 🧵
Advertisement
Advertisement