Check out our new paper!
We present Uno, a discrete diffusion model that accelerates decoding with an AR model while staying (provably) lossless to its distribution.
Uno beats existing discrete diffusion/self-speculative decoding methods, such as Mercury 2, EAGLE-3, and DFlash.
Diffusion LLMs have two limitations relative to AR models:
(1) Lower quality, and
(2) Slower inference at large batch sizes.
We address this "Uno"
> Retains the AR architecture of LLMs
> Each layer has two sets of weights: AR weights and Diffusion weights
> Diffusion
Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters.
- Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting
Super excited to start as an AI Research Intern at @IFM_MBZUAI in Sunnyvale next week (and to get back to the Bay after 7 years)!
Working with @ssahoo_ and lots of other great people, hope to help train some large open-source models!
(If you're in the Bay hit me up!)