Pinned
Best-of-N is a straightforward LLM alignment algorithm: return the highest reward sample of N attempts
👍simple & effective
👎generation throughput decreases by a factor of N.
🤔Can we keep the 👍 while eliminating the 👎? 🧵
w/@ryandcotterell @xtimv
arxiv.org/pdf/2407.06057…

