Multimodal AI Lab

School of Electrical Engineering, KAIST

Announcements

We are looking for motivated students in machine learning, speech processing and computer vision. Please read this page for more information.

Recent highlights

Text-to-Speech


Tan Dat Nguyen et al. (2026), "SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS", Proc. ICASSP

Audio generation from video


Kang Zhang et al. (2025), "Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation", Proc. NeurIPS

Tactile-driven Localization


Seongyu Kim et al. (2026), "Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions", Proc. CVPR

Sign language recognition


Youngjoon Jang et al. (2025), "Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues", Proc. CVPR

Audio-visual LLM decoding

Chaeyoung Jung et al. (2025), "AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding", Proc. NeurIPS

Lip to speech


Ji-Hoon Kim et al. (2025), "From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech", Proc. CVPR


KAIST logo