i keep seeing people ask what the role of grad school and academia is nowadays
but where else do you imagine the production of superstar leaders in a mere few years (or in the case of @a1zhang, a mere few months)
rather than (or at least before!) having them get absorbed into
I'm SUPER EXCITED to publish the 142nd episode of the Weaviate Podcast with Alex Zhang (@a1zhang)! 🔥
Alex is a Ph.D. student at MIT, where he has lead the work behind "Recursive Language Models", as well as "The Mismanaged Genius Hypothesis", "Language Model Harnesses are
god i had almost forgotten how immediate it is to get mouth watering results when doing colbert stuff - @antoine_chaffin, @bclavie, @raphaelsrty, and @aaxsh18 were not lying; i know no other ML area with low-hanging fruit this big
await new open models from @dianetc_ and me soon
Thanks to @tomaarsen for letting me contribute a tiny experiment (the dashed triangles) to this yesterday.
I wanted to demonstrate my claim that late interaction's far stronger quality does NOT generally require a "storage overhead", with technology (PLAID) that was released in
the level to which this was trivial* and yet isn't done by some vector DB providers who offer late interaction by piggybacking on horribly inefficient single-vector infra (then blaming that on the paradigm itself!) is instructive
*i estimate that i hand-held codex for maybe 1hr
Thanks to @tomaarsen for letting me contribute a tiny experiment (the dashed triangles) to this yesterday.
I wanted to demonstrate my claim that late interaction's far stronger quality does NOT generally require a "storage overhead", with technology (PLAID) that was released in
Thanks to @tomaarsen for letting me contribute a tiny experiment (the dashed triangles) to this yesterday.
I wanted to demonstrate my claim that late interaction's far stronger quality does NOT generally require a "storage overhead", with technology (PLAID) that was released in
📈 New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers.
As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find.
Thread 🧵