🐸 We propose a super efficient approach for mechanistic interpretability: decompose weight matrices from a pretrained LLM into sparse circuit units directly, instead of training a separate sparse representation.
See more in the blog post:
huggingface.co/spaces/veri-sa…





