Day 1 with Qwen3.8-2.4T-A95B: 2x Kimi-K3's throughput on one 8x B300 node ⚡
blog.us.fixstars.com/day-1-with-qwe…
11,833 tok/s peak · 48 concurrent requests · 0 failures. 2.4T params on a single node via NVFP4.
Speed up your Embedded AI.
- Agentic AI could fix what's broken about industrial visual inspection. Getting it to run on a factory floor is an optimization problem. Our notes from Dr. Mahbubul Alam's session at the 2026 Embedded Vision Summit: blog.us.fixstars.com/event-report-e…
- People assume AI gets faster if you just speed up the software. But in production, one bottleneck anywhere across the five layers — app to infrastructure — and the whole thing stalls. That's the part people miss. This article is about how LLM training and inference actually
- We're exhibiting at Ai4 2026, booth 1502. Come talk to us about getting AI models to actually hit real-time on edge hardware. We port, optimize, and validate on the target silicon, and we back it with a measurable performance guarantee. Stop by if you're on the floor 👋
- Running a 754B-parameter MoE model on a single H200 node at usable speeds took more than just fitting it in memory. fixstars.com/en/whitepaper/… New white paper from our team on GLM-5.1 inference: where the memory goes, why more GPUs won't necessarily save you, and how P90 TTFT

