Pinned
Ever wondered what makes language models generate overly verbose, vague, or sycophantic responses?
Our new paper investigates these and other idiosyncratic biases in preference models, and presents a simple post-training recipe to mitigate them! Thread below 🧵↓



