Potential crux with MIRI-like rationalists: is it in fact the case that our current world, with Anthropic's influence, is worse than one without Anthropic?
On rationalist views, the world was going get worse and worse anyway (as capabilities advance and we get closer to doom). Anthropic accelerated and continues to accelerate capabilities progress. But how much did they comparatively accelerate alignment and saner AI policy?
In a world with eg. just OpenAI and GDM at the frontier, if/when OpenAI pulls ahead at RSI (as currently seems to be the case):
- would there even have been the current level of integration with UK AISI, current level of model organisms and safety evals?
- would the AI safety space have the expected hundreds of billions of funding, to ambitiously scale its work, including AI policy work?
- would there be have been an AGI company with *some* Operational Adequacy, to proactively do things like Glasswing and biorisk-mitigation? (imo evaluating on the specified criteria, it's clear Anthropic is ahead of OpenAI on most dimensions, and can continue improving on these. One can be upset they aren't technically held by their initial RSP, and yet in practice they seem to be better than OpenAI at it.).
If you wonder why I compare to OpenAI rather than nothing, it's because I don't think "nothing" is the counterfactual of Anthropic not existing. When evaluating the wisdom of Anthropic doing what it did, it's necessary to evaluate against more likely counterfactuals. Possibly many rationalists do take these counterfactuals carefully into account, but the arguments often raised often skip that part. "Anthropic accelerated capabilities" is not a sufficient argument to expect Anthropic's influence on the world to have been net negative.
There are definitely Fabricated Option Worlds which seem much better than the one we got, and on the margin one can hope Anthropic to have done better work or not accelerated capabilities as much, but it' seems difficult fro