Today I'm releasing LlamaStash 0.0.2: a zero-overhead, terminal-native launcher for llama.cpp.
One Rust binary that's a TUI, a CLI, a daemon, and an OpenAI-compatible proxy.
Demo below 馃У
How much local LLM can you run on an AMD Strix Halo with 128GB memory?
I managed to fit DeepSeek v4 Flash 284B and Gemma 4 E2B on GPU, Whisper and Qwen3.5 4B on NPU.
#strixhalo#amd#deepseek#llamastash
LlamaStash v0.0.6 is out 馃
A experimental ds4 backend runs @antirez DeepSeek-V4 GGUFs through DwarfStar (ds4)
Plus: Lemonade on by default, and saved presets that auto-apply.
llamastash.dev#ds4#AI#deepseek
LlamaStash v0.0.5 is out 馃
New: named launch presets. Tune a model's launch knobs once, name them, reuse them, per-model or per-arch. They live in plain config.yaml, so you can hand-edit, comment, and commit them to your dotfiles.
llamastash.dev
LlamaStash v0.0.4 is out 馃
- Auto launch is now the default: llama.cpp's --fit sizes context and GPU offload.
- A browser UI on a stable port
- Anthropic Messages API support.
llamastash.dev