Log inSign up
Awni Hannun
5,061 posts
@awnihannun

Awni Hannun

@awnihannun
ow knee
awnihannun.com
Joined January 2011
356
Following
44.9K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @awnihannun
    Awni Hannun
    @awnihannun
    Sep 3
    $47.9 million mistake not naming the company 🫶 Heart Hands
    @etnshow
    etn.
    @etnshow
    Sep 3
    JUST IN: NVIDIA's $12,930,300,000 acquisition of Hugging Face contains an easter egg. The number 129,303 is the decimal conversion of Unicode point U+1F917. The 🤗 emoji.
    12
  • @awnihannun
    Awni Hannun
    @awnihannun
    Sep 3
    Qwen 27B dense at 105 tok/s output on an m5 max is pretty bonkers. Breaking down the memory wall one brick at a time.
    @dimashvets
    Dima Shvets
    @dimashvets
    Sep 3
    We’re releasing speculative decoding in our inference engine Uzu, starting with Qwen3.6-27B. On Apple M5 Max with 128 GB of unified memory, our Mirai-M stack reaches 105 output tokens/sec entirely on-device - 2.9× faster than the fastest MLX speculative-decoding implementation
    Image
    19
  • @awnihannun
    Awni Hannun
    @awnihannun
    Jul 21
    Interesting that Apple pays Google ~billion per year for Gemini when they can get a better model for free (GLM 5.2 and soon Kimi K3).
    162
  • @awnihannun
    Awni Hannun
    @awnihannun
    Jun 12
    The video from @angeloskath on local agentic AI with MLX is excellent. I also hear it's one of the most viewed videos in WWDC history 👏 Goes through the basics of agentic AI and how to set it all up to run locally in a very approachable and simple way. The demos are excellent
    Image
    11
  • @awnihannun
    Awni Hannun
    @awnihannun
    Jun 9
    It's very cool that Apple shipped a 20B parameter on-device. You can't put 20B parameters in RAM at any reasonable precision. To make it work they are using pretty exotic architecture by today's standards. A small model predicts from the query (or prompt) which experts to load
    Image
    74
Advertisement
Advertisement