Pinned
I'm keen to share our new library for explaining more of a machine learning model's performance more interpretably than existing methods.
This is the work of Dan Braun, Lee Sharkey and Nix Goldowsky-Dill which I helped out with during @MATSprogram:
🧵1/8
Proud to share Apollo Research's first interpretability paper! In collaboration w @JordanTensor!
⤵️
publications.apolloresearch.ai/end_to_end_spa…
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
Our SAEs explain significantly more performance than before!
1/





