Many monitors are trained & evaluated on prompt-elicited hacking trajectories, where models are explicitly asked to exploit the reward signal. But the real test is whether they catch the training-time hacks that naturally emerge during RL training without hacking instructions.
We are a compact and hardcore research team focused on harnessing the power of Multimodal Reasoning. #Google #UCLA #UMD #PennState


