🐸 We propose a super efficient approach for mechanistic interpretability: decompose weight matrices from a pretrained LLM into sparse circuit units directly, instead of training a separate sparse representation.
See more in the blog post:
huggingface.co/spaces/veri-sa…
Providing abundant verification services, which (could) lead to sentient-friendly; Prev. postdoc with @Yoshua_Bengio @Mila_Quebec | PhD @NUSingapore




