Log inSign up
Unsloth AI
775 posts
Unsloth AI profile banner
@UnslothAI

Unsloth AI

@UnslothAI
Run and train models locally with the Unsloth Desktop app. 🦥 github.com/unslothai/unsl…
San Francisco, CA
unsloth.ai
Joined November 2023
478
Following
96.9K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @UnslothAI
    Unsloth AI
    @UnslothAI
    Aug 11
    Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls +
    Image
    00:00
    275
  • @UnslothAI
    Unsloth AI
    @UnslothAI
    Sep 4
    We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction. Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp. Guide: unsloth.ai/docs/models/gl… GGUF: huggingface.co/unsloth/GLM-5.…
    Image
    @Zai_org
    Z.ai
    @Zai_org
    Aug 26
    Image
    Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog:
    59
  • @UnslothAI
    Unsloth AI
    @UnslothAI
    Sep 3
    You can now run Unsloth GGUFs locally in one-click via Hermes! ✨ Qwen3.8-27B, Qwen3.8-Flash, DeepSeek-V4-Flash and more are all supported.
    Image
    @NousResearch
    Nous Research
    @NousResearch
    Sep 3
    Image
    00:17
    Hermes Desktop now sets up local models in one click. It automatically reads your hardware, picks the best model for you, then downloads it and configures the runtime.
    26
  • @UnslothAI
    Unsloth AI
    @UnslothAI
    Sep 2
    Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: huggingface.co/unsloth/Qwen3.… Guide: unsloth.ai/docs/models/qw…
    Image
    @UnslothAI
    Unsloth AI
    @UnslothAI
    Aug 26
    Image
    Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: unsloth.ai/docs/models/qw… GGUF: huggingface.co/unsloth/Qwen3.…
    64
  • @UnslothAI
    Unsloth AI
    @UnslothAI
    Aug 28
    GLM-5.3 can now be run locally! The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.3 is the strongest open model to date. Guide: unsloth.ai/docs/models/gl… GGUF: huggingface.co/unsloth/GLM-5.…
    Image
    @Zai_org
    Z.ai
    @Zai_org
    Aug 28
    Image
    GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize. Weights: huggingface.co/zai-org/GLM-5.3 Tech blog: z.ai/blog/glm-5.3
    101
Advertisement
Advertisement