Llama3 405B is out with impressive performance across the board and 128k ctx length! Along with improved models for 8B and 70B.
Blog : llama.meta.com/llama-download…
Paper : ai.meta.com/research/publi…
A short 🧵 on what went into scaling and training the largest compute Llama model yet
Research Engineer @MetaAI
- It is really amazing how good these models turned out to be!
- In addition to data and model optimization, stability, efficiency and fault tolerance of the entire infra stack is extremely crucial when we scale LLM training to tens of thousands of GPUs. Really glad to be working with @reducescatter and the brilliant AI Infra team @AIatMeta,
- Happy to be part of this incredible journey of Llama3 and to share the best open weight 8B and 70B models! Our largest 400B+ model is still cooking but we are providing a sneak peek into how it is trending! Check more details here
- Excited to be part of the Llama 2 release today and congrats to the entire team for this massive effort! A short thread on my thoughts on the pretrained models, why the results are interesting and where there is room for improvement. 🧵 (1/n) ai.meta.com/llama



