Come chat about @somin's intriguing CoT distillation results — put reasoning after labels, permute or keep only a few key CoT tokens — in Miami!
Joined July 2014
- Sheridan has some cool results on tokenization in LLMs and their "implicit vocabularies"👇(new preprint) LLMs live in a strange tokenized world. We find that LLMs learn to deal with the weirdness of tokenization by converting tokens into word-like representations and then "forgetting about" those tokens. Maybe this is why tokenization isn't an issue, until it is... 🧵
- Somin has some cool results on CoT + distillation👇📢We know that including CoT rationales as supervision improves model distillation. But why? New work (w/ @silvio_amir and @byron_c_wallace) research unveils surprising insights! 🔗 Full paper: arxiv.org/abs/2406.14511 [1/5]
- Learn about "Retrieving Evidence from EHRs with LLMs: Possibilities and Challenges" by @JeredMcinerney, @silvio_amir, @byron_c_wallace at #CHIL2024!
- Come chat with me or @ChantalShaib about @SunJiuding's work on the sensitivity of models to instruction phrasings at #ICLR2024 next week ↓How robust are the instructions in your instruction-tuned model? In our most recent work (w/ @ChantalShaib and @byron_c_wallace), we show that there is a considerable dip in performance on in-domain tasks when you slightly vary the instruction. arxiv.org/abs/2306.11270





