Skip to content
Wei-Cheng (Wayne) Chiu

Wei-Cheng Chiu

邱偉誠

CV

Taipei, Taiwan · UTC+8

making every FLOP count

About

Upstream contributions across the LLM-inference stack span ImageFlashInfer, ImagevLLM, ImageSGLang, ImagePyTorch, ImageDynamo, ImageNVIDIA CUTLASS / TensorRT-LLM, and ImageLMDeploy — with work on correctness, distributed inference, and serving performance. See the auto-updating PR wall and Patches.

Production LLM infrastructure for real customer environments, from site constraints and preflight validation to acceptance and handover. Work spans Kubernetes, GPU runtime, networking, storage, certificates, and model-serving endpoints.

Experience taking an enterprise multi-agent platform from PoC to production, with a focus on model serving, performance evaluation, and solution architecture.

Focus

Broadly, I care about inference performance you can trust — fast kernels that also compute the right answer. Three threads:

  1. Serving internals: KV-cache, quantization trade-offs, attention kernels — enough depth to reason about cost and latency at design time.
  2. Upstream enablement: early consumer-Blackwell (SM120) and NVFP4 support across kernels → engines → disaggregated serving; my favorite prey is the silent-correctness bug — tests green, answers wrong.
  3. Trustworthy ML: safety alignment against harmful fine-tuning; federated learning, differential privacy, secure multi-party computation.

News

Jul 2026
💼 Joined Taiwan AILabs — on-prem LLM deployments (FedGPT).
Jul 2026
🧱 The live PR wall went up — every upstream patch, auto-updating.
Apr 2026
🎓 M.S. in Computer Science from NTUST, GPA 4.09.
Sep 2025
💼 Joined SYNCROBOTIC — sole developer of an enterprise multi-agent platform, shipped at two customers.
Jun 2025
💼 Summer at Advantech building internal coding agents.
Dec 2024
📜 NVIDIA DLI certificates — Accelerated Computing with CUDA (Python & C/C++).
Aug 2024
🔬 Started graduate research on LLM security & privacy-preserving ML at NTUST.

Selected Work

SM120 / NVFP4 enablement across the LLM-inference stack

FlashInfer · CUTLASS · vLLM · SGLang · TensorRT-LLM · Dynamo — kernels to engines to disaggregated serving

CUDABlackwellNVFP4

Live PR wall — prs.wayne.is-a.dev

Auto-updating feed of every upstream contribution, with RSS

Open SourceLive status

PR wall preview

Misc

I'm from Taiwan 🇹🇼 and based in Taipei. Away from a profiler you'll find me tending an over-engineered Obsidian vault. The views on this site are my own and do not represent those of my employer or affiliated institutions.