This is an important problem that IMO is under-addressed because of how difficult it is to measure. Unlike most physical sciences, we can measure things fully and losslessly, but distilling this into measurements that are both interpretable and meaningful is the real challenge.
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵


