Check our new work on cross-modal audio-video generation. Our work produces audio with the best alignment we have seen with respect to actions happening on video. Particularly useful in the era of astounding progress in generative video models.
Can pretrained diffusion models connect for cross-modal generation?
📢 Introducing AV-Link ♾
Bridging unimodal diffusion models in one framework to enable:
📽️ ➡️ 🔊 Video-to-Audio
🔊 ➡️ 📽️ Audio-to-Video
🌐: snap-research.github.io/AVLink/
📄: hf.co/papers/2412.15…
⤵️ Results




