Today we're shipping Nemotron 3 Ultra.
A 550B MoE frontier-intelligence open model built for long-running agents.
It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
Another open-weight release from @thinkymachines 👀 Inkling-Small is here.
With native reasoning over audio and images and variable thinking effort, it's a great choice for fine-tuning, with NVIDIA NeMo on NVIDIA DGX Station.
NVFP4 checkpoint here: huggingface.co/thinkingmachin…
Today, we are releasing Inkling-Small.
Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.
thinkingmachines.ai/news/inkling-s…
Fine-tune it on Tinker today, or chat with
Today we’re open-sourcing Numbat, an agent-detection and response layer that is designed to work across agent harnesses.
Numbat gives security teams visibility into agent activity, with controls to block selected actions before execution.
Read more: research.perplexity.ai/articles/secur…
"Data stops being something you collect. It becomes something you compute."
At #SIGGRAPH2026, @liu_mingyu, VP of NVIDIA's Cosmos Lab, laid out why world models are the data engine of physical AI. Plus, he announced Cosmos-Dreams, a neural closed-loop simulator, and showed off a