Log inSign up
BURKOV
24.4K posts
BURKOV profile banner
@burkov

BURKOV

@burkov
Books: theLMbook.com & theMLbook.com App: ChapterPal.com PhD in AI, author of 📖 The Hundred-Page LMs Book & The Hundred-Page ML Book
Québec, Canada
linktr.ee/burkov
Joined June 2009
127
Following
58.2K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @burkov
    BURKOV
    @burkov
    May 22
    As you know, last year I resigned from my full-time job and become a professional technical book writer. I have plans for The Hundred-Page Books about reinforcement learning, computer vision, diffusion models, and more. Not having a full-time job has put a significant strain on
    Image
    10
  • @burkov
    BURKOV
    @burkov
    2h
    A short book "What Do Mathematicians Do? An Overview of Undergraduate Mathematics" by Steven Clontz and John Estes (2025) is now in @ChapterPal's library. The book provides undergraduate mathematics students with a broad practical guide to mathematical thinking, community
    Image
  • @burkov
    BURKOV
    @burkov
    3h
    The irony is that Attention is All You Need in 2017 meant that you could remove the recurrence from the language model architecture that was considered SOTA at that time (e.g. LSTM with attention) and only keep the attention to get a model that would perform better because of the
    9
  • @burkov
    BURKOV
    @burkov
    5h
    Asked one scientist if I can share his research paper on @ChapterPal. He asked if he will make money from this. I said no. He said "then no." Crazy shit.
    4
  • @burkov
    BURKOV
    @burkov
    5h
    In long-context LLMs, there's a tradeoff: Transformer attention can use the full past but becomes expensive as sequences grow, while recurrent models keep a fixed-size memory whose update rule is itself a form of online learning. This work, from researchers from @ByteDanceSeed_,
    Fast Weight Attention for Continual Learning cover preview
    Fast Weight Attention for Continual Learning — Read with an AI tutor
    From chapterpal.com
    4
Advertisement
Advertisement