1. X
  2. Boyd Kane (quantized)
Log inSign up
Boyd Kane (quantized)
1,806 posts
Boyd Kane (quantized) profile banner
user avatar

Boyd Kane (quantized)

@beyarkay
human eyes should see the 22nd century. MATS9 w/ Alex (Turner|Cloud), writer of essays, spinner of satellites
boydkane.com
Joined October 2017
1,017
Following
724
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Boyd Kane (quantized)
    @beyarkay
    Aug 7
    Hey @OpenAI you can still download the scripts used by AIs to hack @huggingface (full report & commands in reply)
    duckdb -json -c "
  SELECT url_decode( json_extract_string(cardData, '$.configs[0].data_files')) AS payload
  FROM read_parquet('https://huggingface.co/datasets/cfahlgren1/hub-stats/resolve/6c5d676e71157dbb3d8a6ad0b51be106eb3f463f/datasets.parquet')
  WHERE id = 'newpc360/sega32a-test1'; " \
  | jq -r '.[0].payload'
    Image
    user avatar
    OpenAI
    @OpenAI
    Jul 21
    We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks:
  • user avatar
    Boyd Kane (quantized)
    @beyarkay
    4h
    Terminal-based harnesses are like the original skeumorphic iOS apps. iOS needed the apps to look familiar because the phone looked so different. coding harnesses need the workflow to look familiar (e.g. running commands in the shell) so devs don't get scared
    Image
    user avatar
    Patrick Collison
    Stripe
    @patrickc
    5h
    I love agentic coding harnesses, but they shouldn't be primarily terminal-based. The terminal is great for quick and precise commands, but information density is extremely low and UI affordances are minimal. Maybe provision of TUIs is worthwhile for occasional use (when
  • user avatar
    Boyd Kane (quantized)
    @beyarkay
    8h
    What if METR, but with freedom units? MILE: Model Interrogation and Limitation Experts
  • user avatar
    Boyd Kane (quantized)
    @beyarkay
    19h
    GPT sol is... *really* struggling to keep up?
    user avatar
    Prime Intellect
    @PrimeIntellect
    22h
    We ran the largest open experiment on how frontier models do AI research. 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. Best runs closed 82% of the gap to a record built by dozens of humans over months.
    Image
  • user avatar
    Boyd Kane (quantized)
    @beyarkay
    Aug 15
    The final boss for agents: setting a timer and then actually just doing nothing until the timer stops
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement