Tencent releases AuK, a unified speech generation and editing model
It handles zero-shot TTS, instruction-driven content/acoustic/paralinguistic editing, plus enhancement and separation through one natural-language interface. The distilled AuK-Flash version runs 4.5x faster.
Tweeting interesting papers submitted at huggingface.co/papers.
Submit your own at hf.co/papers/submit, and link models/datasets/demos to it!

