deepseek v4.1 flash has arrived on together ai
it beats gpt-5.6 sol on agentic benchmarks at one third the cost per task
the model supports a 1m context window and native multimodal input, with 552b total parameters
together fine-tuning now supports glm-5.3, kimi k2.7-code + more
we also added live run metrics, experiment comparisons, early stopping, dataset previews + finer training controls
prices are down 30–70% on selected models
read more on what's new 👇
introducing preemptible compute for together gpu clusters
same nvidia gpu infrastructure, 50% of the on-demand price
built for evals, fine-tuning, batch inference + short experiments, with up to 5 minutes to checkpoint before reclamation
now in public preview
thunderkittens is now running on nvidia vera rubin!
our kernels team got early access to nvl72 and spent the past few days digging through the new isa and bringing nvfp4 + fp8 gemms to life
after reworking the kernels for rubin, we pushed them past 22 and 12 pflops respectively