1. X
  2. Roman Leventov
Log inSign up
Roman Leventov
3,606 posts
Roman Leventov profile banner
user avatar

Roman Leventov

@leventov
AI engineer. Thinking about hybrid intelligence, AI safety, and AI impacts. [email protected] for contact.
Bali
engineeringideas.substack.com/archive
Joined October 2010
786
Following
1,827
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    Roman Leventov
    @leventov
    Aug 5
    @danluu fuzzing must be great but how do you develop anything about it? It's a bogeyman for codex. As soon as your agent mentions "fuzzing", regardless of the context, you get blocked immediately 😤 @thsottiaux
    Image
    Image
  • user avatar
    Roman Leventov
    @leventov
    Jun 20
    People are worried about fully automated companies, I think for good reasons. What about fully automated non-profits and charities, would that be good or dystopic?
  • user avatar
    Roman Leventov
    @leventov
    May 13
    I think §5.4, building new and adapting old institutions, at large scale, is what a large fraction of OpenAI's foundation and Anthropic's wealth should be directed to. Technical AI safety and biorisk mitigation can absorb maybe up to ~10B/year each, but not ~100B/year.
    user avatar
    Séb Krier
    @sebkrier
    May 12
    If anyone builds it, everyone thrives. Over the past decade, a lot of important work on AI alignment has focused on avoiding harm. But freedom from harm isn't the same as freedom to flourish. In this paper, we introduce 'Positive Alignment'. A positively aligned agent is one
    Image
  • user avatar
    Roman Leventov
    @leventov
    May 8
    This post is more signal about real agent autonomy than the notorious METR chart. I've also run ~30h hour goal with little steering and can corroborate: GPT-5.5-xhigh loses the plot badly. Prioritisation/big picture grasp/engineering common sense are at junior level still.
    user avatar
    Peter Gostev (SF 24-28 August)
    Arena.ai
    @petergostev
    May 5
    I know everything is a skill issue, but I didn't have a good experience with /goal in Codex CLI. I tried giving it 2 tasks, that were hard but well specified. 1) Build a video stabiliser on top of open source to match the Google Photos stabiliser in quality. I had a starter and
    Image
  • user avatar
    Roman Leventov
    @leventov
    Apr 24
    So far, I definitely see a lot of practical engineering judgement regressions between gpt-5.4 and gpt-5.5. I'm not very impressed. The 'leash' couldn't be made longer than it used to be for 5.4
Advertisement
Advertisement