Pinned
Generalist AI engineer.
I created tortoise.cpp, localwriter, and localpilot.
DM's always open.
github.com/balisujohn
- DeepSeek-V4-Flash-0731 (UD-IQ3_XXS) running locally passed the vibe check as a usable automatic Rust programmer. I did not think for a second that something in the ballpark of Opus4.5 would be running at usable speed on my computer in under a year when Opus4.5 came out!
- I found Laguna-S-2.1 (q4_K_M) to be unusable as an automatic Rust programmer with ollama + opencode, but shockingly good for long context writing analysis.
- Seems like an Opus4.5-class automatic programmer at usable toks/sec on a framework desktop. Will have to try setting this up to verify.Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many
- I have achieved inference on int4 glm5.2 at 0.87 tok/s on a single 128gb framework desktop using colibri.


