Everyone: DeepSeek just appeared out of nowhere! 😱
Me:
- DeepSeek Coder in 2023
- MoE in Feb
- Math in Feb
- VL in March
- V2 in May
- Coder V2 in June
- Prover in August
- V2.5 in September
- VL 2 in December
- V3 in December
They've consistently shipped for 1+ years 😁
It's official! After 18 months of work, Hands-On Generative AI with Transformers and Diffusion Models is published and printed! 😱This was one of my most challenging projects and I'm so proud and happy to see it out there.
Check it out!
- Online in O'Reilly
I’m so happy to announce Gemma 3 is out! 🚀
🌏Understands over 140 languages
👀Multimodal with image and video input
🤯LMArena score of 1338!
📏Context window of 128k
Available in AI Studio, Hugging Face, Ollama, Vertex, and your favorite OS tools 🚀Download it today!
I’m so excited to announce Gemma 3n is here! 🎉
🔊Multimodal (text/audio/image/video) understanding
🤯Runs with as little as 2GB of RAM
🏆First model under 10B with @lmarena_ai score of 1300+
Available now on @huggingface, @kaggle, llama.cpp, ai.dev, and more
Introducing Gemma 3 270M 🔥
🤏A tiny model! Just 270 million parameters
🧠 Very strong instruction following
🤖 Fine-tune in just a few minutes, with a large vocabulary to serve as a high-quality foundation
developers.googleblog.com/en/introducing…
Very excited to share some personal news! @johnowhitaker@pcuenq@multimodalart and I are writing a book with @OReilly about generative ML🤗
We'll cover many topics from theory and practical aspects, discuss creative applications, and more!
What topics would you like to see?
BREAKING NEWS
The Royal Swedish Academy of Sciences has decided to award the 2024 #NobelPrize in Literature to the Attention Is All You Need authors.
Their work has made thousands cry, laugh, or rich and made GPUs go brrr
Graphs are **everywhere**, from social media and knowledge systems to molecules and meshes! 🧑🏫
Want to learn about Machine Learning for Graphs? Check out this thread! 🧵
The ML ecosystem in France is on fire🔥 It has amazing talent and resources. Here are 10 facts you might not know:
1. There are great research labs - from @MistralAI and @kyutai_labs to large ones from @AIatMeta and @GoogleDeepMind. The Llama 2 and CodeLlama authors are based in
🧵Stable Diffusion weights are officially public, and we got some surprises! 🤗
🤗 Public weights
🧨Support with the diffusers library
🔥Load and use the model with a few lines of code
📖Blog post explaining how it works
github.com/huggingface/di…
Microsoft just silently dropped Florence
👀Vision model that can tackle many vision tasks (captioning, detection, region proposal, OCR)
🤏Small models (200M and 800M) with ~quality to models 100x larger
🔥MIT licensed
Paper and models: