Log inSign up
Sarah Schwettmann
1,736 posts
Sarah Schwettmann profile banner
@cogconfluence

Sarah Schwettmann

@cogconfluence
Co-founder and Chief Scientist, @TransluceAI, prev @MIT
dessert of the real
cogconfluence.com
Joined October 2015
922
Following
3,220
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @cogconfluence
    Sarah Schwettmann
    @cogconfluence
    Jul 10
    Studying model behavior is both scientifically rich and immediately impactful. It's also difficult to do well, for reasons we discuss in this post! If you're interested in model behaviors, we're working on a lot of ambitious projects, and we're hiring: jobs.gem.com/transluce/am9i…
    @TransluceAI
    Transluce
    @TransluceAI
    Jul 9
    To effectively oversee AI systems, we need to measure how they behave in the world, not just their capabilities. In a new essay, we describe our vision for an open scientific ecosystem for model behavior evaluation, and the public infrastructure required to support it.
    Image
  • @cogconfluence
    Sarah Schwettmann
    @cogconfluence
    Dec 18, 2025
    All @TransluceAI work that I described in my NeurIPS mech interp workshop keynote is now out! ✨ Today we released Predictive Concept Decoders, led by @vvhuang_ Paper: arxiv.org/pdf/2512.15712 Blog: transluce.org/pcd And here's @damichoi95's work on scalably extracting
    @JustinAngel
    Justin Angel
    @JustinAngel
    Dec 7, 2025
    We can train models on maximizing how well they explain LLMs to humans 🤯@cogconfluence paraphrased. Mechanistic Interpretability Workshop #NeurIPS2025.
    Image
    Image
    1
  • @cogconfluence
    Sarah Schwettmann
    @cogconfluence
    Nov 25, 2025
    My favorite part of @damichoi95’s new paper (alongside 2 new datasets!) is the scaled up investigator pipeline that directly decodes open-ended user representations from model internals end-to-end interp is increasingly promising and I'm excited for more work in this direction
    Image
    @TransluceAI
    Transluce
    @TransluceAI
    Nov 25, 2025
    What do AI assistants think about you, and how does this shape their answers? Because assistants are trained to optimize human feedback, how they model users drives issues like sycophancy, reward hacking, and bias. We provide data + methods to extract & steer these user models.
  • @cogconfluence
    Sarah Schwettmann
    @cogconfluence
    Nov 25, 2025
    Come say hi at #NeurIPS2025! @TransluceAI is hosting a lunch event on Thursday where we'll discuss our recent work on understanding AI systems and where we're headed next. Would love to see you there 👇
    @TransluceAI
    Transluce
    @TransluceAI
    Nov 24, 2025
    Transluce is headed to #NeurIPS2025! ✈️ Interested in understanding model behavior at scale? Join us for lunch on Thursday 12/4 to learn more about our work and meet members of the team: luma.com/8kjfb378
    1
  • @cogconfluence
    Sarah Schwettmann
    @cogconfluence
    Aug 31, 2025
    found two things in the de Young sculpture garden today that I had no idea were here! a beehive piece from Pierre Huyghe, who I worked with in 2022 to install a hive in simulation (along with a real one) on an island in Norway…
    Image
    Image
    Image
    1
Advertisement
Advertisement