Pinned
Still amazed by the talent in this community.
<2 months after GTC, we rebuilt the stack end-to-end — kernel → scheduler → modeling → frontend — and shipped the fastest open-source MLA attention on Blackwell for agentic workloads.
Grateful to be part of it!
Introducing TokenSpeed, a speed-of-light LLM inference engine.
> TensorRT LLM level performance
> vLLM level usability
> Built by a lean and mission-driven team in two months
> MIT license, open-source
github.com/lightseekorg/t…
lightseek.org/blog/lightseek…





