Pinned
Code for "Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types" is now available!
It includes B-TAP, the interp method behind our work, for localizing and causally intervening on critical weights.
Code + project page + paper 👇🏻



