Pinned
@iclr_conf paper alert! The de facto way to align a model through tuning-based methods like DPO is powerful, yet expensive and prone to jailbreaking. Emerging work on model editing aims to address this, and yet the two approaches are largely siloed. Can we somehow connect them?🧐



