Log inSign up
Hadas Orgad
319 posts
@OrgadHadas

Hadas Orgad

@OrgadHadas
Research Fellow @ Kempner Institute, Harvard | Working on AI interpretability, robustness & safety
orgadhadas.github.io
Joined April 2019
151
Following
1,257
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @OrgadHadas
    Hadas Orgad
    @OrgadHadas
    Aug 26
    Code for "Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types" is now available! It includes B-TAP, the interp method behind our work, for localizing and causally intervening on critical weights. Code + project page + paper 👇🏻
    Image
    Image
    1
  • @OrgadHadas
    Hadas Orgad
    @OrgadHadas
    Sep 1
    Come work with me! I’ll be mentoring in this fall’s @cbai_ai fellowship. If you’re interested in interpretability + AI safety, especially understanding the mechanisms behind sycophancy, deceptions, harmful generations and hallucinations, apply to my stream 👇
    @cbai_ai
    Cambridge Boston Alignment Initiative
    @cbai_ai
    Aug 31
    Applications are open for the CBAI Fall Research Fellowship in AI Safety & AIxBiosecurity. Apply by September 6th at 11:59 PM EDT! 📅October 13 - December 18 💵 $15,000 stipend for 10 weeks 💰 ~$1,500+/week compute support 🏠 Housing close to our offices, arranged by CBAI
    Image
    6
  • @OrgadHadas
    Hadas Orgad
    @OrgadHadas
    Aug 24
    Announcing the keynote speakers for the Actionable Interpretability Workshop at COLM 2026! @ActInterp @boknilev @banburismus_ @ChrisGPotts @dhanya_sridhar What would you ask them?
    Image
    2
  • @OrgadHadas
    Hadas Orgad
    @OrgadHadas
    Aug 22
    Accepted to #emnlp ✅ 🎉 Read about the (non-) robustness of uncertainty probes 👇🏼
    @_joestacey_
    Joe Stacey
    @_joestacey_
    Apr 15
    Replying to @_joestacey_
    This work has been really fun to work on! Massive thank you my fantastic collaborators @OrgadHadas @inuikentaro @benbenhh and @NafiseSadat You can find the paper here: arxiv.org/abs/2604.11662
  • @OrgadHadas
    Hadas Orgad
    @OrgadHadas
    Jul 26
    Accepted to COLM 2026? Consider submitting your work to the fast track at the Actionable Interpretability Workshop! @ActInterp We welcome work that advances the use of interpretability to real-world usage. Deadline: August 9 Links below 👇
    Image
    1
Advertisement
Advertisement