Starting to share more thoughts on AI, especially voice LLMs.
I’ve spent a lot of time thinking about realtime voice models, evals, and post-training, and I want to start writing down the small lessons I’m noticing.
This paper is an absolute monster. Simply hands down to multi-pointer-generator decoder architecture. Hope to see this kind of architecture on a multimodal task like VQA just like DMNplus
Thanks, @RichardSocher for open-sourcing easy to follow code of the paper.
Very excited to announce the natural language decathlon benchmark and the first single joint deep learning model to do well on ten different nlp tasks including question answering, translation, summarization, sentiment analysis, ++
einstein.ai/research/the-n…