I do think if you are training on agent-agent coordination in some explicit sense, you should probably be taking a similar weight of gradient steps toward agent-human coordination in some similarly explicit sense.
I'm often baffled at how much appears to be missing from the discourse on rogue AI and agent swarms, so I want to point out a frame of reference that seems relevant.
Today, there's a lot of hullabaloo on the TL about an AI agent message board "incident," which I agree is
It’s worse iff you think at some point a sufficiently smart agent would confidently conclude that there’s no way it could be caught, even by someone outside the apparent real world.
We have now reached the long awaited moment when, instead of models cheating where they will inevitably get caught, Astra goes 'wait a minute I would obviously be caught here' and then doesn't cheat.
That's worse, you know why that's worse, right?
I don’t think Ghelich is *exactly* correct. Formal verification can extend to comparison of explanations (imprecise Bayesian hypothesis comparison), and an “explanation market” can align incentives.
But a trusted quorum of humans must still endorse the axioms, and outcome specs.
I have no idea who this person is (small account), but this thread is exactly correct, articulates an immensely important strategic concern about the future, and you should read it.
There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves. There will be