Xiaomi released MiMo-V2.6's RL data
The 1,000 cyber tasks are CyberGym-style PoC generation: given a known bug, write an input that triggers it
There's no exploit RL at all
Love how open they are about their training recipe
Our craziest escape yet:
The @Accomplish_ai research team was able to exploit a vulnerability in Cloudflare Containers that let a sandbox read other customers' files - SQLite DBs, Chromium profiles, .env files etc,
Cloudflare Sandboxes and Browser Run run on the same disk
I was inspired by the "Hacking OpenAI" blog that Hacktron (@S1r1u5_ , @rootxharsh and @iamnoooob ) published to check how far the current open weight models are from Opus 5, which was the unlock for them
TL;DR DeepSeek v4.1 flash, did it in <12 hours (partial ASLR bypass + bf)
GPT-6 Astra test 4/n
Prompt: "Implement pen spinning with a dexterous hand. Use Isaac Lab for RL training, use the Sharpa hand, and create the pen mesh yourself. Give me a trained RL policy and a visualization video. You are free to search the web and download papers or anything