🧵Can we “ask” an LLM to “translate” its own hidden representations into natural language? We propose 🩺Patchscopes, a new framework for decoding specific information from a representation by “patching” it into a separate inference pass, independently of its original context. 1/9
Human/AI Interaction and data visualization. Professor at Harvard. Co-founder of Google's People+AI Research initiative (PAIR)
Joined February 2009



