|
|
B3-PWL: GPU-Batched Branch-and-Bound for Piecewise-Linear Optimization with SOS2 Constraints
Yilin Guan,
Shuqing Luo,
Pingzhi Li,
Tianlong Chen,
Kaidi Xu
arXiv preprint, 2026
arXiv
A GPU-centric branch-and-bound framework for piecewise-linear optimization with SOS2 constraints: batches of LP relaxations are solved concurrently with a first-order primal-dual solver and a batched block-tiled sparse kernel, with a unified feasibility-search module for fast incumbents — 9.25x geometric-mean speedup over NVIDIA cuOpt on 43 PWL-MIP instances.
|
|
|
AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
Shuqing Luo*, Yilin Guan*, Pingzhi Li, Hanrui Wang, Tianlong Chen
* Equal contribution
ICML, 2026
arXiv
Asynchronous framework for efficient test-time scaling: light-weight temporal-regressive query prediction and disaggregated KV-cache filtering overlapped with inference.
|
|
|
Dynamic Speculative Agent Planning
Yilin Guan,
Wenyue Hua,
Qingfeng Lan,
Fei Sun,
Dujian Ding,
Devang Acharya,
Chi Wang,
William Yang Wang
ICLR, 2026
arXiv
/
code
A dynamic speculative planning framework for LLM-based agents that accelerates multi-step reasoning by adaptively speculating future actions.
|
|