Log inSign up
Tomasz Limisiewicz
310 posts
@TomLimi

Tomasz Limisiewicz

@TomLimi
Postdoctoral researcher at @meta Fair and @uwnlp , Interested in going into the inner workings of neural networks, multilingualism, and fairer NLP (he/him)
Seattle
Joined September 2021
517
Following
836
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @TomLimi
    Tomasz Limisiewicz
    @TomLimi
    May 4
    We present Compute Optimal Tokenization! 🔡 Common in LLM scaling works stick to one tokenizer, sweeping data/model size. But what happens when we control the tokenizer’s compression rate (bytes/token)? Here we sweep tokenizers, params, and data across compute budgets: [1/N]
    Image
    00:00
    21
  • @TomLimi
    Tomasz Limisiewicz
    @TomLimi
    Sep 4
    Solve that Astra…
    @zouharvi
    Vilém Zouhar
    @zouharvi
    Sep 4
    Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
    3
  • @TomLimi
    Tomasz Limisiewicz
    @TomLimi
    Sep 1
    If you are still skeptical about byte LMs and think they are just too slow, check out Abraham’s new paper! Showing that employing multi-token (👉byte👈) prediction speeds up generation without sacrificing great performance!
    @AbrahamOwos
    Abraham Owodunni
    @AbrahamOwos
    Aug 27
    Our new paper is online: Dynamic Multi-Byte Prediction With Hierarchical Language Models Hierarchical byte-level LMs have begun to gain increasing adoption due to their tokenizer-free advantage, but they lag behind subword models with respect to inference speed. 🧵 1/n
    Image
    00:00
    1
  • @TomLimi
    Tomasz Limisiewicz
    @TomLimi
    Aug 27
    From now on, I vibecode only in Polish 🇵🇱
    @magikarp_tokens
    Sander Land
    @magikarp_tokens
    Aug 26
    Looked a bit more at Claude's language focus. Slavic is an even bigger outlier: 25% of v4.7's vocabulary, double any other tokenizer (as a share of its small vocab, not raw slots). Meanwhile the spread for Indic is also huge! Explore yourself @ tokenize.rs/compare
    Image
    1
  • @TomLimi
    Tomasz Limisiewicz
    @TomLimi
    Jul 29
    small vocab is the future
    @magikarp_tokens
    Sander Land
    @magikarp_tokens
    Jul 29
    Got quite a bit further with figuring out Claude's tokenizer, and finally wrote it all up. Enjoy! open.substack.com/pub/tokencontr…
Advertisement
Advertisement