Most ARM chips can't run decent AI models. Introducing EfficientSpeech, a 266k-param TTS model. Low cost ARM chips like in RPi4 can generate 104sec of speech mel spec in 1sec. Here's an AI-generated video w/ voice from EfficientSpeech. Info: github.com/roatienza/effiโฆ #ICASSP2023
Creator of ViTSTR and EfficientSpeech (ICASSP2023) and co-creator of PARSeq. Professor & Scientist at the University of the Philippines.
- We are lucky to have this paper accepted at ECCV 2022! More details later ... #ECCV2022Scene Text Recognition with Permuted Autoregressive Sequence Models deepai.org/publication/scโฆ by Darwin Bautista and @jacobe #Computation #Language
- Idea: If data augmentation improves model generalization, why not use it to generate 2 new inputs and force the representations to agree. Result: Additional model performance improvement. Comparison: Unlike Label Smoothing, the performance of our method, AgMax, is consistent.Improving Model Generalization by Agreement of Learned Representations from Data Augmentation deepai.org/publication/imโฆ by @jacobe #ComputerVision #ImageNet
- Yesterday, my former grad student Daryl gave a talk at Sony CSL Paris about his thesis on Next View Policy for 3D Reconstruction. Youtube: youtu.be/KdyDj3bjU0I


