I am a postdoctoral researcher (and Young Investigator) at the Allen Institute for AI and the University of Washington, advised by Prof. Hanna Hajishirzi. Additionally, I am part-time affiliated with the ETH AI Center, where I mentor students and work on post-training for the Swiss AI Initiative.
I completed my PhD in Computer Science at the NLP lab of Bar Ilan University, supervised by Prof. Ido Dagan and Prof. Reut Tsarfaty. I also was a visiting PhD student at UW NLP, with Prof. Yejin Choi, and had the pleasure of interning twice at the Allen Institute for AI. My work has been awarded an ACL Outstanding Paper Award and the ACL Best Theme Paper Award. I am also very honored to have received the AI2 Outstanding Intern of the Year Award.
Previously I did a research internship at Google, obtained an MSc from the University of Edinburgh and a BA from the University of Zurich. My work has been featured in the press, for example by TechCrunch and GeekWire.
Work with me: Swiss AI Initiative
I am looking for motivated students to work with me on contributing to the Swiss AI Initiative’s LLM efforts, particularly with a focus on post-training. We are open to host students at ETHZ or EPFL for semester or thesis projects.
Indicate interest using this application form, mentioning my name.
AI+X Summit: Staging × Misalignment Science
Join us for the Staging × Misalignment Science workshop at the AI+X Summit in Zürich: when, where and how safety enters LLM training — from data curation and pre-training to post-training, evaluation, interpretability, and deployment.
Learn more and register: AI+X Summit session page
Research
My research focuses on how one can develop generative AI that is contextually robust, responsible and open. In particular, I have focused on extending language models’ capabilities through post-training and adaptation. Additionally, I have been involved in the construction of multiple, widely-used benchmarks, such as RewardBench! More specifically, my research is centered around:
- Open Science of LLMs and Post-Training: Developing good open recipes for language model post-training.
- I am a core contributor on the Tulu and Open-Instruct project, where we develop post-training pipelines consisting of supervised finetuning, direct preference optimization, and reinforcement learning with verifiable rewards.
- I have worked on the open science of language models, by contributing to OLMo and OLMo 2.
- Steerability, Underspecification and Context: Improving how models deal with contextual robustness, underspecified inputs, and how they can respond more precisely to instructions.
- Critical Evaluation: Building challenging benchmarks for more realistic, human-centered evaluation of generative models and reward models.
Awards
- Aug 2024ACL Outstanding Paper Award!
- Aug 2024ACL Best Theme Paper Award for OLMo!
- Oct 2023Was selected as a DAAD AInet fellow.
- Feb 2023Was awarded a postdoctoral scholarship from the Eric and Wendy Schmidt Foundation.
- Jan 2023Was awarded the AI2 Outstanding Intern of the Year Award.
- Jan 2021Awarded the Nadav Award for Excellence in Research.
Publications
Below is a selection of my recent publications; for my full publication record, please see my Google Scholar page.
2026
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
RewardBench 2: Advancing Reward Model Evaluation
2025
Olmo 3
IF-RLVR: Generalizing Verifiable Instruction Following
TÜLU 3: Pushing Frontiers in Open Language Model Post-Training
2 OLMo 2 Furious
Diverging Preferences: When do Annotators Disagree and do Models Know?
SafetyAnalyst: Interpretable, transparent, and steerable LLM safety moderation
RewardBench: Evaluating Reward Models for Language Modeling
Superlatives in Context: Modeling the Implicit Semantics of Superlatives
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
2024
Explicating the Implicit: Argument Detection Beyond Sentence Boundaries
Self-Directed Synthetic Dialogues and Revisions Technical Report
The Art of Saying No: Contextual Noncompliance in Language Models
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
OLMo: Accelerating the Science of Language Models
Promptly Predicting Structures: The Return of Inference
Retrieving Texts based on Abstract Descriptions
2023
Camels in a Changing Climate: Enhancing LM Adaptation with TÜLU 2
"You Are An Expert Linguistic Annotator": Limits of LLMs as Analyzers of Abstract Meaning Representation
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
Design Choices for Crowdsourcing Implicit Discourse Relations: Revealing the Biases Introduced by Task Design
ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral Situations
Revisiting Sentence Union Generation as a Testbed for Text Consolidation
2022
Just-DREAM-about-it: Figurative Language Understanding with DREAM-FLUTE
QASem Parsing: Text-to-text Modeling of QA-based Semantics
Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and Training
Draw Me a Flower: Grounding Formal Abstract Structures Stated in Informal Natural Language
2021
Asking It All: Generating Contextualized Questions for any Semantic Role
The Possible, the Plausible, and the Desirable: Event-Based Modality Detection for Language Processing
2020
QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines
QA-Nom: Question-Answer driven SRL for Nominalizations
2017
Discourse Relations and Conjoined VPs: Automated Sense Recognition
* denotes equal contribution.
Misc
Besides this I love rowing, hiking, and going to the “cinemathèque”. I think that Italian Neorealism produced some of the most beautiful movies. My Erdős number is 3 (Paul Erdős → Noga Alon → Ido Dagan → Me) and my Kevin Knight number is 2 (Kevin Knight → Yejin Choi → Me).