Pinned
Training an LLM on 8 M4 Mac Minis
Ethernet interconnect between Macs is 100x slower than NVLink so Macs can’t synchronise model gradients every training step.
I got DiLoCo running so Macs synchronise once every 1000 training steps using 1000x less communication than DDP

