SkyRL + Harbor RL recipe!
Worked with Mercor Research and the SkyRL team, training Qwen3.5-397B-A17B on APEX-Agents (long-horizon office work) off-the-shelf data with SkyRL, improving Pass@1 from 16% to 27%.
The post is more of a practical field guide for what to do given an RL dataset, de-risking step






