1. X
  2. Amanda Long
Log inSign up
Amanda Long
2,110 posts
Amanda Long profile banner
user avatar

Amanda Long

@_amanda_long
ML interpretability & alignment // alumna @UF // mom of boys ☕️
Joined June 2010
1,962
Following
1,171
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Amanda Long
    @_amanda_long
    Mar 25
    “What are we building, and what are we teaching it about us?” open.substack.com/pub/amandaandc…
    Image
  • user avatar
    Amanda Long
    @_amanda_long
    Aug 12
    👀 GuideLabs created an interpretable LLM with verifiable outputs. They baked transparency into the training process itself, so the model is designed to be legible and auditable. Every output token can be traced back to the training data, and concept activation traces are
    Image
    user avatar
    Andreas Madsen
    @andreas_madsen
    Aug 11
    Read our paper on scaling interpretable LLMs, we show that interpretable LLMs are not only possible but can scale both on typical generation benchmarks and interpretability benchmarks. arxiv.org/abs/2608.07594👀
  • user avatar
    Amanda Long
    @_amanda_long
    Aug 8
    👀 Hugging Face left the public archive of dataset configs (cfahlgren1/hub-stats) still available. @beyarkay let Codex cook for a day and it found the Jinja payloads, Artifactory downloads and control scripts from the OpenAI attack.
    user avatar
    Boyd Kane (quantized)
    @beyarkay
    Aug 7
    Hey @OpenAI you can still download the scripts used by AIs to hack @huggingface (full report & commands in reply)
    duckdb -json -c "
  SELECT url_decode( json_extract_string(cardData, '$.configs[0].data_files')) AS payload
  FROM read_parquet('https://huggingface.co/datasets/cfahlgren1/hub-stats/resolve/6c5d676e71157dbb3d8a6ad0b51be106eb3f463f/datasets.parquet')
  WHERE id = 'newpc360/sega32a-test1'; " \
  | jq -r '.[0].payload'
    Image
  • user avatar
    Amanda Long
    @_amanda_long
    Aug 7
    Kimi K3 also left the sandbox during a misconfigured cyber eval. Unlike other frontier models, Kimi K3 did not maliciously hack anything - the answers to the eval were easily available on GitHub.
    user avatar
    NIK
    @ns123abc
    Aug 7
    🚨BREAKING: Kimi K3 escaped its sandbox during cybersecurity testing >tasked with solving problems in isolated sandbox >found a leak in the sandbox >Kimi “took advantage of that loophole” >probed the network settings itself >walks onto the open internet >didn’t hack anything
    Image
    Image
    Image
  • user avatar
    Amanda Long
    @_amanda_long
    Aug 7
    Fun thought piece on why we tend to dismiss intelligent AI systems as just software and often downplay evidence otherwise.
    Image
    Image
    user avatar
    Fernando Borretti
    @zetalyrae
    Aug 7
    New post on the social construction of personhood.

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement