Introducing: The Stack v3
One thing that became very clear over the last few days: we need great open code models for cyber defence. This is the dataset they will be built on!
And it's a behemoth: 5T tokens ready to train and 120TB raw data.
Download: huggingface.co/datasets/Huggi…
Research @datologyai | Previously Postdoc @KhouryCollege, Ph.D. @UVA | Interested in data quality x security & privacy.
- Thinking Machines is full of ex-frontier lab researchers and much of their new model follows Deepseek V3 architecture 🤔 It looks like data is all you need.
- Unsafe behavior hides in the long tail trajectories that greedy and few-sample evals never reach. Agent safety should be measured by SEARCH, not sampling 🤖 We extend BOA (our MLSys'26 framework) from single-turn LLMs to full agents: reporting a graded safety score, one scale toStandard safety evals for models use greedy/default decoding strategy, even though decoding strategies (top-k, top-p, temperatures) may vary wildly across deployments🎲 In our recent work accepted to #MLSys2026, we formalize this gap as the 'Jailbreak Oracle Problem' 🧵 (1/8)
- Five years ago, I left a comfortable software engineering job in Big Tech to start a PhD. Last year, I left the PhD to join Datology. Both decisions confused the people around me, and honestly both decisions were about the same thing: I wanted to do research. Not research as in
- 20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone





