Log inSign up
Lucas
639 posts
Lucas profile banner
@quantbagel

Lucas

@quantbagel
inference @hf0, ex-quant 2x
SF | Montreal
Joined August 2025
2,507
Following
1,629
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @quantbagel
    Lucas
    @quantbagel
    Mar 3
    Robot action models shouldn't need 256 vision tokens per frame. Pi0.5 spends 400M parameters on SigLIP just to see. We replaced it with a 4.4M encoder that outputs 5 tokens — and action quality barely changes. 91x smaller. 51x fewer tokens. 7.3x faster inference.
    Image
    26
  • @quantbagel
    Lucas
    @quantbagel
    Sep 8
    Async heterogenous inference is going to be one of the major inference markets soon. You can disaggregate the models across a bunch of hardware, prefill on a cluster of Macs, dflash on a gb200, the kvcache on nvme.
    Image
    15
  • @quantbagel
    Lucas
    @quantbagel
    Sep 8
    is there a harness built specifically for arbitrary perf optimization? Accross arbitrary chip types as well
  • @quantbagel
    Lucas
    @quantbagel
    Sep 6
    theres a hidden model in codex at the moment, ask astra to find it! Its a long context version of astra (I think)
    Image
  • @quantbagel
    Lucas
    @quantbagel
    Sep 4
    Anyone need a VM breaking RL harness
Advertisement
Advertisement