1. X
  2. Marc Sun
Log inSign up
Marc Sun
721 posts
user avatar
Marc Sun
@_marcsun
Machine Learning Engineer @huggingface Open Source team
New york
Joined February 2023
508
Following
1,576
Followers
RepliesRepliesMediaMedia
  • user avatar
    Marc Sun
    @_marcsun
    Jan 29
    If you could fix ONE thing about `Trainer` in transformers, what would it be? Share your feedback: github.com/huggingface/tr… Thanks @UnslothAI @axolotl_ai and others for trusting and building on top of Trainer. We want to make sure you all get the best experience.
    Image
    Tell Us: What Would Make Trainer Better? · Issue #43595 · huggingface/transformers
    From github.com
  • user avatar
    Marc Sun
    @_marcsun
    Dec 1, 2025
    Amazing day 🤩
    user avatar
    Lysandre
    @LysandreJik
    Dec 1, 2025
    Transformers v5's first release candidate is out 🔥 The biggest release of my life. It's been five years since the last major (v4). From 20 architectures to 400, 20k daily downloads to 3 million. The release is huge, w/ tokenization (no slow tokenizers!), modeling & processing.
    Image
  • user avatar
    Marc Sun
    @_marcsun
    Aug 28, 2025
    Happy to participate in the online course by my mentor @TheZachMueller ! The topic of my talk will be efficient distributed inference
    user avatar
    Zach Mueller
    Lambda
    @TheZachMueller
    Aug 28, 2025
    14 Days of Distributed, Day 12! Meet Marc Sun (@_marcsun) of @huggingface Marc is a ML Engineer working on the Open Source team at Hugging Face and he collaborates with researchers and developers to add new exciting features in the HF ecosystem and have contributed to various
    Image
  • user avatar
    Marc Sun
    @_marcsun
    Aug 8, 2025
    We’ve made N-D parallelism training simpler in accelerate ! It was a pleasure to collaborate with @axolotl_ai on this feature.
    user avatar
    Matej Sirovatka
    Prime Intellect
    @m_sirovatka
    Aug 8, 2025
    We have cooked some something nice with @axolotl_ai for 🤗 accelerate v1.10. Have you ever wanted to train a large model, but couldn't setup your env? Current multi-gpu frameworks are hard to install, the configuration has like 100 options. Here comes ParallelismConfig 🚀 1/5 🧵
    Image
  • user avatar
    Marc Sun
    @_marcsun
    Aug 5, 2025
    Great post from @mekkcyber about mxfp4 format that was used in the gpt-oss models !
    user avatar
    Mohamed
    White Circle
    @mekkcyber
    Aug 5, 2025
    The new GPT-OSS models are Mixture of Experts (MoEs), with 20B and 120B parameters. Since expert weights make up ~90% of the model, OpenAI decided to quantize them to 4 bits during post-training using the MXFP4 standard. Quantizing these to MXFP4 enables the larger model to
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement