New blog post with @sarahwiegreffe on OpenAI's Monitorability Evals!
We hope to make others working with these evals aware of some weaknesses we came across, and encourage more work on chain-of-thought monitorability evals.
Excited to announce my first preprint in LM interpretability!
Latent reasoning models are not monitorable by default, since they don't reason in human-readable, natural language text. But can we make progress in understanding their intermediate reasoning steps using mech interp?
Excited to announce my first preprint in LM interpretability!
Latent reasoning models are not monitorable by default, since they don't reason in human-readable, natural language text. But can we make progress in understanding their intermediate reasoning steps using mech interp?