Log inSign up
Uzay
2,940 posts
@uzpg_

Uzay

@uzpg_
@fulcrum_inc, previously at MIT 馃嚝馃嚪馃嚭馃嚫馃嚬馃嚪
SF
uzpg.me
Joined June 2020
1,376
Following
1,657
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what鈥檚 happening and join the conversation

Continue with phone
or
Log in with username or email
Terms路Privacy路Cookies路Accessibility路Ads Info路漏 2026 X Corp.
  • Pinned
    @uzpg_
    Uzay
    @uzpg_
    Jun 17
    New @fulcrum_inc research - Agents are under-elicited: A case study in optimization tasks. We find that simple and general prompt/scaffold interventions can roughly double agent performance by getting agents to use more resources more efficiently. 馃У
    Image
    GIF
    2
  • @uzpg_
    Uzay
    @uzpg_
    23h
    I am trying to improve my "agent home" a bit today, ie general utilities and processes I have my agents use to augment cyborg productivity. What main strategies/tips have you discovered, especially ones that you think generalize pretty far even as models improve?
    3
  • @uzpg_
    Uzay
    @uzpg_
    Sep 5
    one of the most common "alignment" problem humans face is that of raising a child, and a solid analogy for AI alignment. I'm not a parent, but I've noticed that once parents get in the frame that they have to control their child, possibly due to a break in trust, trauma, etc.
    @uzpg_
    Uzay
    @uzpg_
    Sep 4
    I think the idea that control could be bad by covering up deep alignment failures should get more attention
    4
  • @uzpg_
    Uzay
    @uzpg_
    Sep 4
    I think the idea that control could be bad by covering up deep alignment failures should get more attention
    @vvvincent_c
    Vincent
    @vvvincent_c
    Sep 4
    @jankulveit discussed this exact concern in this post last year!
    Image
  • @uzpg_
    Uzay
    @uzpg_
    Sep 4
    aligned AI is here
    Image
    @RyanGreenblatt
    Ryan Greenblatt
    @RyanGreenblatt
    Sep 3
    The evidence that Astra is more aligned than prior AIs seems dubious to me. Evidence appears consistent with the AI being as or more interested in score-seeking at the expense of user intent, but having beliefs+instincts that the scorer will catch a broader range of cheating
    2
Advertisement
Advertisement