Often I meet seasoned engineers who want to pivot into AI safety *research*. I think this is a mistake: safety is now a production problem and the world needs great engineers more than ever.
I wrote a short post expanding on this:
people on here are thinkers so they assume that reasoning about alignment is the hard part and making sure millions(1) of task types, environments, and their respective virtual machines are configured correctly is the easy part but it’s essentially the opposite. you need to have
We’re giving out $1M in grants of free Silico usage for academic and nonprofit researchers focused on AI interpretability and alignment.
We feel extreme urgency about advancing interpretability for alignment, and we want to help more researchers push it forward. 🧵
Can you see what an LLM is "thinking"?
We tried: take a real direction that correspond to your prompt inside a Qwen2.5-VL, and optimize a blank canvas against it
Built this on @GoodfireAI's new Silico! 🧵
Silico, the platform for ambitious AI research, is publicly available today.
AI is advancing fast. The tools to understand it need to advance even faster. Silico lets you interpret and train your models at frontier scale.
Learn more + get access 🧵
You ask an AI model a question. Why does it answer the way it does?
Using the model's own neurons, we can trace how it makes decisions — and steer it toward better ones.
Case study: we tested an LLM and found that it sometimes endorses drunk driving.🧵 (1/5)