If you could fix ONE thing about `Trainer` in transformers, what would it be?
Share your feedback: github.com/huggingface/tr…
Thanks @UnslothAI @axolotl_ai and others for trusting and building on top of Trainer. We want to make sure you all get the best experience.
New york
Joined February 2023
- Amazing day 🤩Transformers v5's first release candidate is out 🔥 The biggest release of my life. It's been five years since the last major (v4). From 20 architectures to 400, 20k daily downloads to 3 million. The release is huge, w/ tokenization (no slow tokenizers!), modeling & processing.
- Happy to participate in the online course by my mentor @TheZachMueller ! The topic of my talk will be efficient distributed inference14 Days of Distributed, Day 12! Meet Marc Sun (@_marcsun) of @huggingface Marc is a ML Engineer working on the Open Source team at Hugging Face and he collaborates with researchers and developers to add new exciting features in the HF ecosystem and have contributed to various
- We’ve made N-D parallelism training simpler in accelerate ! It was a pleasure to collaborate with @axolotl_ai on this feature.We have cooked some something nice with @axolotl_ai for 🤗 accelerate v1.10. Have you ever wanted to train a large model, but couldn't setup your env? Current multi-gpu frameworks are hard to install, the configuration has like 100 options. Here comes ParallelismConfig 🚀 1/5 🧵
- The new GPT-OSS models are Mixture of Experts (MoEs), with 20B and 120B parameters. Since expert weights make up ~90% of the model, OpenAI decided to quantize them to 4 bits during post-training using the MXFP4 standard. Quantizing these to MXFP4 enables the larger model to








