1. X
  2. Jacob Andreas
Log inSign up
Jacob Andreas
2,855 posts
Jacob Andreas profile banner
user avatar

Jacob Andreas

@jacobandreas
Teaching computers to read. Assoc. prof @MITEECS / @MIT_CSAIL / @NLP_MIT (he/him). lingo.csail.mit.edu web.mit.edu/jda/www
Cambridge, MA
Joined March 2007
956
Following
24.7K
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Jacob Andreas
    @jacobandreas
    Jul 3
    👉 New preprint (we had a big backlog 😅)! Revisiting adversarial imitation learning for the era of RLVR:
    user avatar
    Mehul Damani
    @MehulDamani2
    Jul 3
    Higher benchmark scores do not always mean better models for users. Why? We claim that RL teaches LMs to be correct but not how to be correct: code can pass tests but be unreadable; explanations can be right but unclear. How do we train LMs to be right in the right way? (1/n)
    Image
  • user avatar
    Jacob Andreas
    @jacobandreas
    Jul 24
    👉 New preprint! (and a direction I'm excited to do more work on soon)
    user avatar
    Elinor
    @elinorpd_
    Jul 23
    Pluralistic alignment is thriving as a research agenda yet failing at its goal: making the AI systems people actually use more pluralistic🌈 🚨New position paper: we argue adoption in deployed models should be the fields main goal & we provide a roadmap of how to get there 🧵1/
    Image
  • user avatar
    Jacob Andreas
    @jacobandreas
    Jul 7
    If you're excited about Anthropic's J-space work, definitely worth checking out the original paper on Jacobian lenses by @evanqed and @arnab_api!
    arXiv logo
    arxiv.org
    Linearity of Relation Decoding in Transformer Language Models
    Much of the knowledge encoded in transformer language models (LMs) may be expressed in terms of relations: relations between words and their synonyms, entities and their attributes, etc. We show...
  • user avatar
    Jacob Andreas
    @jacobandreas
    Jul 1
    👉 Preprint: understanding learning dynamics & mechanisms in LMs trained to explain / predict their own behaviors!
    user avatar
    Carl Guo
    @CarlGuo866
    Jul 1
    New Paper 📄: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of
    Image
  • user avatar
    Jacob Andreas
    @jacobandreas
    Jun 29
    👉 New preprint! Automated interpretability by approximating / replacing NN components (here attention heads) with programs.
    user avatar
    Amiri Hayes
    @amirihayes_
    Jun 29
    What if attention were code? We show that many attention heads in transformer LMs can be replaced by human-readable Python programs. Swap them in and the model barely notices. See our experiments here: Explaining Attention with Program Synthesis [arxiv.org/abs/2606.19317]
Advertisement
Advertisement