Log inSign up
Sebastian Raschka
20.2K posts
Sebastian Raschka profile banner
@rasbt

Sebastian Raschka

@rasbt
ML/AI research engineer. Ex stats professor. Author of "Build a Large Language Model From Scratch" (amzn.to/4fqvn0D) & reasoning (mng.bz/lZ5B)
United States
sebastianraschka.com
Joined October 2012
1,204
Following
505.5K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @rasbt
    Sebastian Raschka
    @rasbt
    Jul 18
    How can an LLM switch between low-, medium-, and high-effort reasoning? And how does an LLM learn to reason more or less? I put together a “little” article explaining how these effort levels are implemented at inference time and during training.
    Image
    114
  • @rasbt
    Sebastian Raschka
    @rasbt
    13h
    I put together a mega write-up on GPT-6 Astra & looped transformers. How looped transformers / recurrent depth works, cost-tradeoffs, whether it hides reasoning traces, with lots of figures and a tour of recent looped transformer research.
    Image
    48
  • @rasbt
    Sebastian Raschka
    @rasbt
    Sep 8
    Re today's incident, maybe not a bad idea to check your settings (Settings → Data Controls)
    Image
    91
  • @rasbt
    Sebastian Raschka
    @rasbt
    Sep 6
    Reasoning from scratch round 2: In this video, I cover the text generation process in LLMs and KV caching (to prepare the base model before adding reasoning techniques in the upcoming ones). 00:00 Introduction and reasoning model demo 01:55 How to work through the book 05:00
    Image
    00:00
    48
  • @rasbt
    Sebastian Raschka
    @rasbt
    Sep 2
    A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth or looped transformer". It's always interesting to read about new or different approaches (including
    Image
    137
Advertisement
Advertisement