"alignment at the individual level" is useful but doesn't necessarily result in coordination. if we want AIs that coordinate sensibly, we need to create both learners & learning environments specifically for coordination
anthropic.com/research/multi…
Softmax's mission is to scale organic alignment. We approach this problem with multi-agent reinforcement learning population-based simulations.
- it turns out that if you have fun in a game, you have it in real life too
- It’s Annealing Week at Softmax! Humans are awake for 16 hours learning, cooling for 4 hours in light sleep, and in deep sleep for 4. An organic mental annealing cycle, heating to cooling. At Softmax, we do the same. It’s four weeks sprinting towards goals, one week consolidating.
- Our little Cogs grow up so fast. Cogbert has never seen this exact production chain before, but with only a couple missteps he begins to execute it correctly. Our in-context learner takes its first baby steps!
- We are building organic alignment at Softmax. Not just with reinforcement learning, but within our company we try to use these same principles for our work. We are implementing this as an organizational operations system (OrgOS), a prompt library covering our internal processes.

