One thing that I’m wondering is what sort of benchmarks (with easy evaluation) are out there that stress-test the same type of tasks humans pose to chatGPT. Any pointers?
After all the great research coming up in the community around the pitfalls of static LMs, these news from @huggingface are particularly exciting! There is a great deal of applications that online LMs can enable and a lot of interesting questions around new evaluations!
We’re going to do it! We’ll train and release masked and causal language models (e.g. BERT & GPT-2) on new Common Crawl snapshots as they come out! We call this project Online Language Modeling (OLM). What applications or research questions can we enable or help answer? A 🧵:
Less than a week left for the paper submission of the inaugural EMNLP Industry Track! Paper submission June 25th. Looking forward to the exciting submissions!
EMNLP Call for paper for the inaugural *industry track* is posted! 2022.emnlp.org/calls/industry…
We have three tracks — deployed, emerging, and discovery, all focusing on real-world implementations of NLP systems. The submission deadline is July 25, 2022.
If you are at ICML check out our work on keeping QA models updated! Some reasons of why I’m super excited about this work! First, we make creative use of few-shot prompting to create a large QA dataset grounded in different points in time. Synthetic data FTW! 1/N