Three years ago we showed seq-to-seq models could be used for text-to-image synthesis. In some way I’m glad to see new models are emerging in 2021/2022 where this idea is scaled up and combined with visual dictionary learning (ie. VQ-VAE, VQ-GAN, dGAN).
Look Mom No GANs: Text-to-image Synthesis as Machine Translation // Not the actual title of our paper but maybe should be. To be presented next week at #cvpr2019 in Long Beach, California.



