Log inSign up
Sebastian Raschka
20.1K posts
Sebastian Raschka profile banner
@rasbt

Sebastian Raschka

@rasbt
ML/AI research engineer. Ex stats professor. Author of "Build a Large Language Model From Scratch" (amzn.to/4fqvn0D) & reasoning (mng.bz/lZ5B)
United States
sebastianraschka.com
Joined October 2012
1,191
Following
498.8K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @rasbt
    Sebastian Raschka
    @rasbt
    Jul 18
    How can an LLM switch between low-, medium-, and high-effort reasoning? And how does an LLM learn to reason more or less? I put together a “little” article explaining how these effort levels are implemented at inference time and during training.
    Image
    109
  • @rasbt
    Sebastian Raschka
    @rasbt
    Aug 30
    A little video that - explains the relationship between conventional LLMs and reasoning models (and agents), - philosophizes a about "from scratch" approaches, - and explains how to install Python & PyTorch requirements with uv.
    Image
    00:00
    40
  • @rasbt
    Sebastian Raschka
    @rasbt
    Aug 29
    Thanks everyone for all the nice feedback on Build a Reasoning Model (From Scratch) so far! I am also flattered that 2 book clubs are discussing it. I look forward to join for live Q&As on Thu, Sep 3, at 10 am (and 2 pm CT). Please join us & bring your questions!
    @sophiamyang
    Sophia Yang
    @sophiamyang
    Aug 27
    Join The AI Book Club for a live conversation with @rasbt about his new book, Build a Reasoning Model (From Scratch)! 📅 Sep. 3, 10 AM CT 🌐 Online We'll talk about how reasoning models work, how to build them, and take questions from the community.
    Image
    18
  • @rasbt
    Sebastian Raschka
    @rasbt
    Aug 26
    Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash... Compared to GLM-5.2, this new GLM-5.3-Flash model uses: - a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Sparse
    Image
    Image
    @Zai_org
    Z.ai
    @Zai_org
    Aug 26
    Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog:
    56
  • @rasbt
    Sebastian Raschka
    @rasbt
    Aug 22
    A couple of days ago, I did a quick explainer on Claude’s new watermarking process and implementation. Since it’s such a popular topic and sparked such a lively discussion, I thought it might be interesting to go into a bit more detail when explaining how it works. So, instead
    Image
    39
Advertisement
Advertisement