explore VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- Speed-of-light block sparse attention :🚀 Sol-Engine weekly update is here! This week we released some new things: • Integrated Sol Attention — our team’s new training-free sparse attention method for video diffusion (full paper coming next week 🚀). It’s already available in the engine as an early easter egg. •
- Congrats!Today, SkyPilot is out of stealth. Building custom intelligence is now existential. We help frontier AI teams build intelligence faster by removing their biggest bottleneck: AI compute fragmentation. Frontier teams like @appliedcompute, @AbridgeHQ, @hippocraticai, @hcompany_ai,
- Explore KDA (Kernel Design Agents): github.com/mit-han-lab/ke…Databricks ranks #1 on NVIDIA’s SOL-ExecBench kernel leaderboard, in the L1 single operation track, powered by KDA (Kernel Design Agents) 🎉 What’s crazy is: we 100% leveraged AI agents to beat the competition. This is a sneak peek at recursive self-improvement. The core
- We develop an agent-native approach to accelerate genAI, continuing the success of KDA (Kernel Design Agent) at a higher level. We distill each technique into a skill: step distillation, sparse attention, token pruning, quantization, kernel fusion, then let agents integrate them🚀 Sol Video Inference Engine is here! An agent-native, training-free full-stack accelerator for video diffusion. It auto-tunes cache + sparse attn + token pruning + quant + kernel fusion for any model/hardware/config. >2× end-to-end speedup on 64B Cosmos3-Super, 22B LTX-2.3





