- The most exciting breakthroughs in intelligence are yet to come. I’m super excited to start this journey with mes amis to make them happen together.
- I was debating with @ducx_du in the past few days on a few points, sharing them to provide some more food for discussion. There can be two interpretations, 1. DLLM is fitting a loose elbo with uniform posterior distribution over the order. 2. It is fitting n! number of modelsDiffusion LLMs (DLLM) can do “any-order” generation, in principle, more flexible than left-to-right (L2R) LLM. Our main finding is uncomfortable: ➡️ In real language, this flexibility backfires: DLLMs become worse probabilistic models than the L2R / R2L AR LMs. This






