Good points by Daniel, worth reading his post!
new post! I argue that current alignment techniques might soon become obsolete (as we scale RL), and furthermore might obfuscate misalignment.
This draws on evidence from the recent Anthropic / OpenAI cybersecurity incidents, among other things.
lesswrong.com/posts/nLaQmJf4…





