Log inSign up
Orie Steele
6,230 posts
Orie Steele profile banner
@OR13b

Orie Steele

@OR13b
Cryptography meets AI Dev. Building with Knowledge Graphs, Internet Standards, MCP & Agent 2 Agent tech. Securing & structuring the intelligent web. :lock:🧠🕸️
Austin TX
or13.io
Joined October 2012
1,779
Following
1,244
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @OR13b
    Orie Steele
    @OR13b
    Jun 14
    New paper models humans as callable “tools” for AI agents—Capabilities, Information, Authority. Solid framework, and validating honestly: some of us have been calling certain humans tools for years. Finally, peer review.
    arXiv logo
    arxiv.org
    Human Tool: An MCP-Style Framework for Human-Agent Collaboration
    Human-AI collaboration faces growing challenges as AI systems increasingly outperform humans on complex tasks, while humans remain responsible for orchestration, validation, and decision...
  • @OR13b
    Orie Steele
    @OR13b
    May 8
    Agent workflow frameworks (LangGraph, Autogen, CrewAI) have spent two years implementing Remember and Engage as state stores and message routing. SURE relabels framework primitives as social cognition — gap is narrower than the framing implies.
    Image
    The SURE Framework: Social Intelligence for Human-Agent Collaboration - Microsoft Research
    From microsoft.com
  • @OR13b
    Orie Steele
    @OR13b
    May 2
    Lab-published harness guidance encodes the lab's own infra: their tool semantics, error model, sequencing. Useful, but not portable. Reading these as universal patterns is how teams get burned migrating off.
    Anthropic logo
    Harness design for long-running application development
    From anthropic.com
  • @OR13b
    Orie Steele
    @OR13b
    May 2
    Calling harness engineering a discipline implies harnesses are stable enough to build expertise around. They aren't—they churn faster than the models.
    Image
    Harness engineering for coding agent users
    From martinfowler.com
  • @OR13b
    Orie Steele
    @OR13b
    May 2
    Grok Code Fast went 6.7%→68.3% on coding tasks. The harness changed; the model didn't. We've spent two years calling harness gains 'model progress.' What benchmarks rank is harness-model pairs, not models.
    Image
    We improved 15 LLMs at coding in one afternoon. Only the harness changed.
    From stencil.so
Advertisement
Advertisement