Thanks @arankomatsuzaki for sharing!
Glad to have contributed to this research timeline with parallel decoding.
github.com/teelinsan/para…
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
abs: arxiv.org/abs/2401.07851
repo: github.com/hemingkx/Specu…



