Log inSign up
Pratyush Maini
DatologyAI
859 posts
Pratyush Maini profile banner
@pratyushmaini

Pratyush Maini

DatologyAI
@pratyushmaini
Assistant Professor @cornell_tech & Founding Team @datologyai | PhD @mldcmu | BTech @iitdelhi
pratyushmaini.github.io
Joined November 2019
618
Following
3,727
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @pratyushmaini
    Pratyush Maini
    DatologyAI
    @pratyushmaini
    Mar 19
    If I had to compress my PhD into one idea, it is this "The data a model sees early in training leaves an imprint on its representations that is very hard to undo later" This thread runs through - Rephrasing the Web - Safety Pretraining - TOFU This is the Finetuner’s Fallacy🧵
    Image
    00:00
    21
  • @pratyushmaini
    Pratyush Maini
    DatologyAI
    @pratyushmaini
    Sep 1
    If you lead your vertical, you have years of data that exists nowhere on the internet. That data is now your most valuable asset and the way to realize that value is to train your own models on it @datologyai x @thomsonreuters is our first public example of what that looks like
    @levie
    Aaron Levie
    Box
    @levie
    Sep 1
    Now that the base open weights AI models are getting far better, and post training infra is becoming more mature and commercialized, there are going to be all new plays for companies that have large amounts of data to have their own models. Licensing data for external model
    2
  • @pratyushmaini
    Pratyush Maini
    DatologyAI
    @pratyushmaini
    Aug 27
    It’s incredible how the surge of money into AI made everyone feel poorer. At the start of the decade, researchers made less, but slept better. Now everyone makes more, compares more, and feels they’re getting less. Somewhere along the way, money acquired satisfaction?
    2
  • @pratyushmaini
    Pratyush Maini
    DatologyAI
    @pratyushmaini
    Aug 18
    The best data researchers are obsessive about looking at data. But doing that at scale takes incredible time and care. This is exactly where swarms of agents can tirelessly sift through petabytes of data across data curation pipelines. People should look at autonomous data
    @oliveraochongli
    Oliver Li
    @oliveraochongli
    Aug 18
    1/ Finally get to share what I’ve been working on this summer at @datologyai. We built DataSmith 🍊, an automated data researcher. The idea is simple: scale test-time compute across multiple directions, then let agents synthesize their findings. Blog: datologyai.com/blog/datasmith
    Image
    00:00
    3
  • @pratyushmaini
    Pratyush Maini
    DatologyAI
    @pratyushmaini
    Aug 18
    true story
    Image
    @HaoliYin
    Haoli Yin
    DatologyAI
    @HaoliYin
    Aug 18
    Image
    General coding agents often lack research taste: they make small tweaks and get stuck in local optima. We built DataSmith to automate data research. It inspects failures, searches wider, and makes better research bets, beating Claude Code by 5.1 pts on avg. 1/n
    1
Advertisement
Advertisement