Log inSign up
Jordan Taylor
541 posts
Jordan Taylor profile banner
@JordanTensor

Jordan Taylor

@JordanTensor
Working on new methods for understanding machine learning systems and entangled quantum systems.
Brisbane
sites.google.com/view/jordanten…
Joined December 2009
1,174
Following
489
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @JordanTensor
    Jordan Taylor
    @JordanTensor
    May 17, 2024
    I'm keen to share our new library for explaining more of a machine learning model's performance more interpretably than existing methods. This is the work of Dan Braun, Lee Sharkey and Nix Goldowsky-Dill which I helped out with during @MATSprogram: 🧵1/8
    Image
    @leedsharkey
    Lee Sharkey
    @leedsharkey
    May 17, 2024
    Proud to share Apollo Research's first interpretability paper! In collaboration w @JordanTensor! ⤵️ publications.apolloresearch.ai/end_to_end_spa… Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning Our SAEs explain significantly more performance than before! 1/
    1
  • @JordanTensor
    Jordan Taylor
    @JordanTensor
    Sep 3
    Unfortunate news for monitorability. Some excerpts from the system card 🧵
    @JBloomAus
    Joseph Bloom
    @JBloomAus
    Sep 3
    Replying to @JBloomAus
    We found Astra can solve much more difficult problems in a single forward pass than past models, with a no-reasoning math time horizon of 30 minutes (!) compared to 4 minutes for GPT 5.6 Sol.
    Image
    1
  • @JordanTensor
    Jordan Taylor
    @JordanTensor
    May 21
    There are a lot of pathways via which AI oversight is likely to degrade! Latent reasoning architectures, situational awareness, representational drift... We wrote a report ranking them. Here I'll go into some which worry me most 🧵
    @AISecurityInst
    AI Security Institute (AISI)
    @AISecurityInst
    May 21
    The safety of advanced AI systems increasingly depends on the ability to oversee them. Our new report examines today’s AI oversight landscape, finding many pathways likely to lead to its degradation.🧵
    Image
    1
  • @JordanTensor
    Jordan Taylor
    @JordanTensor
    May 1
    Come work with me!
    @JBloomAus
    Joseph Bloom
    @JBloomAus
    May 1
    (My team) Model Transparency at @AISecurityInst is hiring Research Engineers and Research Scientists! Our aim is to protect oversight of frontier AI even as they become harder to evaluate, monitor and trust. As capabilities scale, this is becoming a harder and more important
  • @JordanTensor
    Jordan Taylor
    @JordanTensor
    Mar 30
    Excited to see this!
    @satvikgolechha
    7vik
    @satvikgolechha
    Mar 30
    Research from Model Transparency @ UK AISI: we reproduce the Anthropic work "Natural Emergent Misalignment from Reward Hacking in Production RL" using OS models, RL environments, algorithms, and tooling + we share an unexpected result related to CoT faithfulness. 🧵 (1 of 7)
    Image
Advertisement
Advertisement