Log inSign up
Harbor Framework
148 posts
Harbor Framework profile banner
@harborframework

Harbor Framework

@harborframework
San Francisco, CA
harborframework.com
Joined January 2026
3
Following
2,131
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @harborframework
    Harbor Framework
    @harborframework
    Sep 1
    SkyRL + Harbor RL recipe!
    @charlie_ruan
    Charlie Ruan
    @charlie_ruan
    Sep 1
    Worked with Mercor Research and the SkyRL team, training Qwen3.5-397B-A17B on APEX-Agents (long-horizon office work) off-the-shelf data with SkyRL, improving Pass@1 from 16% to 27%. The post is more of a practical field guide for what to do given an RL dataset, de-risking step
    Image
  • @harborframework
    Harbor Framework
    @harborframework
    Aug 29
    there's a new Terminal-Bench in town
    @ryan_marten
    Ryan Marten
    @ryan_marten
    Aug 29
    We've pushed a version update to the Terminal-Bench dataset and leaderboard. Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
    Image
    00:00
    1
  • @harborframework
    Harbor Framework
    @harborframework
    Aug 28
    Braintrust 🤝 Harbor
    @braintrust
    Braintrust
    @braintrust
    Aug 24
    Some agents need a sandbox where they can do work, liking editing files, installing dependencies, and launching builds. To score these agents you run the tests and check what changed, which takes a clean container per attempt. Harbor is a Python framework for specifying
    Image
    1
  • @harborframework
    Harbor Framework
    @harborframework
    Aug 27
    Terminal-Bench-Science is here! Check out terminal-bench-science.ai to see the tasks, results, and trajectories
    @StevenDillmann
    Steven Dillmann
    @StevenDillmann
    Aug 27
    We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions
    Image
  • @harborframework
    Harbor Framework
    @harborframework
    Aug 25
    Evaluate the singularity with @harborframework
    @mhrezaeics
    MohammadHossein Rezaei
    @mhrezaeics
    Aug 25
    We've released our initial tasks and environments in @harborframework format at github.com/scaleapi/rsi-b… Huge thanks to @nas_mahmoud_, @ChenguangWang, and @_yunzhong, as well as @MingchenZhuge and @tydsh (who contributed in their personal time), for their collaboration and
Advertisement
Advertisement