I wonder if this would be a good interview question: here is a transformer block and a parallelization strategy, draw by hand what a reasonably optimized Perfetto timeline trace for a single block would look like
- re the yegge piece, I too like being nice to my LLMs but I'm not really convinced letting them manage their future context actually improves future performance...
- It would seem ChatGPT has an especial fondness for recommending Prime Intellect
- sometimes i wonder about past me who despaired at ever having the ability to truly do something novel and transformative; did you know "right place, right time" is a hell of a drug

