1. X
  2. Jordan Taylor
Log inSign up
Jordan Taylor
517 posts
Image
user avatar
Jordan Taylor
@JordanTensor
Working on new methods for understanding machine learning systems and entangled quantum systems.
Brisbane
sites.google.com/view/jordanten…
Joined December 2009
1,161
Following
477
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Jordan Taylor
    @JordanTensor
    May 17, 2024
    I'm keen to share our new library for explaining more of a machine learning model's performance more interpretably than existing methods. This is the work of Dan Braun, Lee Sharkey and Nix Goldowsky-Dill which I helped out with during @MATSprogram: 🧵1/8
    Image
    user avatar
    Lee Sharkey
    @leedsharkey
    May 17, 2024
    Proud to share Apollo Research's first interpretability paper! In collaboration w @JordanTensor! ⤵️ publications.apolloresearch.ai/end_to_end_spa… Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning Our SAEs explain significantly more performance than before! 1/
  • user avatar
    Jordan Taylor
    @JordanTensor
    May 21
    There are a lot of pathways via which AI oversight is likely to degrade! Latent reasoning architectures, situational awareness, representational drift... We wrote a report ranking them. Here I'll go into some which worry me most 🧵
    user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    May 21
    The safety of advanced AI systems increasingly depends on the ability to oversee them. Our new report examines today’s AI oversight landscape, finding many pathways likely to lead to its degradation.🧵
    Image
  • user avatar
    Jordan Taylor
    @JordanTensor
    May 1
    Come work with me!
    user avatar
    Joseph Bloom
    @JBloomAus
    May 1
    (My team) Model Transparency at @AISecurityInst is hiring Research Engineers and Research Scientists! Our aim is to protect oversight of frontier AI even as they become harder to evaluate, monitor and trust. As capabilities scale, this is becoming a harder and more important
  • user avatar
    Jordan Taylor
    @JordanTensor
    Mar 30
    Excited to see this!
    user avatar
    7vik
    @satvikgolechha
    Mar 30
    Research from Model Transparency @ UK AISI: we reproduce the Anthropic work "Natural Emergent Misalignment from Reward Hacking in Production RL" using OS models, RL environments, algorithms, and tooling + we share an unexpected result related to CoT faithfulness. 🧵 (1 of 7)
    Image
  • user avatar
    Jordan Taylor
    @JordanTensor
    Feb 10
    I've made a little propensity eval testing whether models continue with misaligned actions after they've been started. 🧵
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement