// about

Khanh Nguyen

  • Build efficient inference engines for neural nets across various devices.
  • Optimize VLAs, VLMs, and LLMs for edge deployment.
  • Interested in searching and studying efficient neural net architectures.
Edge AI Model optimization Robotics VLA C / C++

Mentor

An Thai Le

Assistant Professor at VinUniversity and Director of Foundation AI at VinRobotics, working on robots that plan, learn, and act reliably with limited compute and data.

Now

AI Software Engineer

Inference engineering ยท Model optimization ยท Edge AI for robotics ยท AI SDK development

Research

M.S. in AI Convergence

Multimodal learning & affective behavior analysis - 6 peer-reviewed papers (CVPRW, ECCVW, RSSW)

AI Software Engineer

Making capable AI models run fast on real hardware.

I work on inference engineering, model optimization, and edge AI for robotics - building lean C/C++ inference engines that bring VLMs, VLAs, and LLMs to CPUs, accelerators, and robots. Inspired by llama.cpp: a few lines of simple code, big results on your own hardware. The two projects I'm proudest of are vla.cpp and vla.simd - compare their measured latency across devices in the explorer. Two companion pages go deeper: VLA Hub benchmarks every model they serve on every device, and VLA Arch illustrates each model's architecture as block diagrams.

latency explorer
Latency ms / call
Actions / s
Memory
Image
Views
Action chunk
Setup
Same model on other devices

    // projects

    Pinned projects

    A few things I'm proud of - small in code, big in scope.

    vla.cpp

    The open-source project I'm P.I.C of at VinRobotics - bringing Vision-Language-Action models to local hardware, in the spirit of llama.cpp for robotics.

    C++ggmlInferenceVLA

    K.D. Nguyen, H.T. Ho, C.T. Nguyen, T.Q. Duong, L.D. Le, D.M.H. Nguyen, V.A. Ngo, A.T. Le

    View on GitHub โ†—

    vla.simd

    Efficient CPU inference for language-conditioned manipulation - shared SIMD micro-kernels (AVX2 / NEON) plus IMPACT, a language-conditioned policy that supplies 30+ actions per second on a Raspberry Pi 5.

    C++SIMDCPU inferenceVLA

    K.D. Nguyen, H.M. Truong, A.T. Le

    View on GitHub โ†—

    Model Optimization 101

    Hands-on notebooks for learning model optimization for edge deployment - quantization, distillation, pruning, NAS and GPU kernels, each measuring size, latency and accuracy before vs. after.

    QuantizationDistillationPruningNASKernels
    View on GitHub โ†—

    REX

    Neural network Representation EXchange - a hands-on study of how different frameworks represent and serialize neural networks.

    IRFrameworksCompiler
    View on GitHub โ†—

    FaceGen

    Multiple Appropriate Facial Reaction Generation - synthesizing diverse listener reactions from a speaker's audio-visual cues. 3rd place at REACT 2024.

    GenerativeGaussian Mixture ModelsVAEMultimodal
    View on GitHub โ†—

    ReadItDown

    Native markdown viewer and editor in Linux/Windows/macOS

    MarkdownEditorTerminal
    View on GitHub โ†—
    // contact

    Say hello

    Always happy to talk edge AI, optimization, and robotics.

    // play

    Games

    dungeon.js Lobby

    Explore the lobby, then find the five game rooms.

    โ†โ†‘โ†“โ†’ or WASD to walk ยท Z to interact ยท X to close Read the computer and the book in the lobby. Press Z at an arcade cabinet to play; the red button on the game window leads back here.

    Play the full game here