Check out this Mixture of Experts project by our member!!
This week, I wanted to see if we could get the smallest possible Mixture of Experts going that takes a few hours to pretrain on a single GPU, is fast for experiments, but also somehow performs well compared to larger MoEs. For reference, we have Nanochat for transformer




