- a mathematical take on building startups open.substack.com/pub/andylawk/p…
- Something clicked today for me. Interpretability is at odds with scaling. Learning polysemantic features is an efficient solution via SGD. Sparse, monosemantic features cost capacity and gradient flow — it’s foolish to expect these traits to emerge natively


