1. X
  2. Emmanuel Ameisen
Log inSign up
Emmanuel Ameisen
2,179 posts
@mlpowered

Emmanuel Ameisen

@mlpowered
Interpretability/Finetuning @AnthropicAI Previously: Staff ML Engineer @stripe, Wrote BMLPA by @OReillyMedia, Head of AI at @InsightFellows, ML @Zipcar
San Francisco, CA
mlpowered.com/book/
Joined June 2017
247
Following
11.3K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @mlpowered
    Emmanuel Ameisen
    @mlpowered
    Mar 27, 2025
    We've made progress in our quest to understand how Claude and models like it think! The paper has many fun and surprising case studies, that anyone who is interested in LLMs would enjoy. Check out the video below for an example
    @AnthropicAI
    Anthropic
    @AnthropicAI
    Mar 27, 2025
    New Anthropic research: Tracing the thoughts of a large language model. We built a "microscope" to inspect what happens inside AI models and use it to understand Claude’s (often complex and surprising) internal mechanisms.
    Image
    00:00
  • @mlpowered
    Emmanuel Ameisen
    @mlpowered
    Aug 21
    Neural networks compress information very densely. That makes them confusing to interpret. In fact, if you just look at the largest weights in a network, they often do not make sense! This writeup shows why, and explores approaches to identify which weights are useful.
    @nicholasturner0
    Nicholas Turner
    @nicholasturner0
    Aug 21
    Even if we have ways to break up neural networks into interpretable parts, the largest weights between those parts can be confusing. Why is that? In our new research note, we study the *weight* superposition that causes interference weights.
    Image
  • @mlpowered
    Emmanuel Ameisen
    @mlpowered
    Aug 10
    shot chaser
    Image
    Image
  • @mlpowered
    Emmanuel Ameisen
    @mlpowered
    Jul 15
    LLMs can keep track of the length of various pieces of text perfectly. How do they do this? In this new paper, we find that they learn a general mechanism: countdown heads! The mechanism is both simple and useful, and it shows up in a surprisingly large set of examples 🧵
    Image
  • @mlpowered
    Emmanuel Ameisen
    @mlpowered
    Jul 6
    We locate a small subset of the representations inside a language model (the J-space), and find that it resembles an internal monologue. It shows intermediate concepts, as well as judgements about the current text. It even detects prompt injections!
    Image
Advertisement
Advertisement