I audited FineWeb-Edu, the educational subset of FineWeb, for misinformation. I annotated 200K documents from the 100BT sample using Llama 4 Maverick.
Assistant Professor at IT University of Copenhagen
- Happy to share our new paper! We study Tabular Language Models (TLMs): LLMs fine-tuned on serialized tabular data, converting rows into text sequences. Do TLMs actually learn tabular reasoning? Our re-evaluation suggests maybe not. 🧵 Paper:
- Excited to share our latest work on initializing expanded vocabulary for language models! Kudos to Nandini and Aditya for their excellent contributions!🚨🚨 New preprint 🚨🚨 Presenting: An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models Paper: arxiv.org/abs/2407.05841 Code: github.com/AI4Bharat/Voca… @anoopk @ratishsp @nandi_mundra
- Absolutely elated to be a part of this huge effort!
- Multiple acceptances in @emnlpmeeting thanks to students, researchers and collaborators. 1. CTQScorer: Combining Multiple Features for In-context Example Selection for Machine Translation Authors: @NameIsAshwanth, @ratishsp, @prajdabre1, @anoopk (Findings) #EMNLP2023 #NLProc



