Heading to #ICML2026. Excited to reconnect, learn, share what we've been working on (around self-distillation, test-time training, exploration for discovery etc.) and experience Seoul!
Turns out self-distillation (via SDPO) is also a promising way to align, personalize and continually adapt LLMs directly from user conversations, without requiring any explicit preference labels or reward models! @ETH_AI_Center
Deployed LLMs and users generate millions of conversations every day.
These are full of useful learning signals, yet we don't use them for training.
We introduce self-distillation for learning directly from user conversations – no rewards, no labels, no extra models.
SDPO enables RL agents to learn from rich feedback (i.e., not only whether an attempt failed, but why it failed, such as error messages). Even without such rich feedback, SDPO can reflect on past attempts and outperform GRPO. SDPO accelerates solution discovery at test time!
Training LLMs with verifiable rewards uses 1bit signal per generated response. This hides why the model failed.
Today, we introduce a simple algorithm that enables the model to learn from any rich feedback!
And then turns it into dense supervision.
(1/n)
Exciting work by our startup @latticeflowai ! Assessing risks associated with AI models on a deep technical level is essential for their successful deployment!
@ETH_AI_Center
🚨 We have BIG news to share today, we’re introducing AI Insights - the First Independent LLM Evaluation Service for Risk & Readiness.
🔗 insights.latticeflow.ai
Reflecting back on an inspiring week at #AIHouseDavos - such a vibrant environment to discuss, with AI thought leaders across academia, industry and the public sector, challenges around safe and responsible AI, and harnessing AI for sustainable development! It was truly inspiring