Many monitors are trained & evaluated on prompt-elicited hacking trajectories, where models are explicitly asked to exploit the reward signal. But the real test is whether they catch the training-time hacks that naturally emerge during RL training without hacking instructions.
We are a compact and hardcore research team focused on harnessing the power of Multimodal Reasoning. #Google #UCLA #UMD #PennState
- We’re excited to share our new work on CVPR 2026, Understanding Reward Hacking in Text-to-Image Reinforcement Learning. Reinforcement learning is becoming an increasingly important tool for post-training text-to-image generation models. But as we optimize these models with
- Many thanks to @_akhaliq for sharing our work on the first multimodal Aha moment with a 2B non-SFT model. Join us on our journey into multimodal reasoning and stay tuned for more cool research at @TurningPointAI turningpoint-ai.com!VisualThinker-R1-Zero R1-Zero's Aha Moment on just a 2B non-SFT Model VisualThinker-R1-Zero is a replication of DeepSeek-R1-Zero in visual reasoning. Successfully observe the emergent “aha moment” and increased response length in visual reasoning on just a 2B non-SFT models
- 🚀 We’re excited to share our latest work! Welcome to the first successful "aha moment" on multimodal reasoning. "Aha moment" is featured by improved response length & performance. It emerges during RL of an unaligned base model on multimodal tasks. Aha moment for language
- We made Multimodal LLMs safe, but have they also become oversensitive? "Every time I try, it uses all tokens just refusing." - @artilectium "This isn’t safety. It's a nanny state." - @krishnanrohit Concerned AI safety has gone too far? you’re not alone! Explore MOSSBench by


