Log inSign up
DatologyAI
259 posts
DatologyAI profile banner
@datologyai

DatologyAI

@datologyai
DatologyAI builds tools to automatically select and optimize the best data on which to train AI models, leading to better, smaller models which train faster.
Redwood City, CA
datologyai.com
Joined September 2023
10
Following
3,338
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @datologyai
    DatologyAI
    @datologyai
    40m
    We sat down with the research teams from Thomson Reuters and DatologyAI to break down how they built the Thomson-1 model. We get into the data, the mid-training, and what it takes to build a frontier model on open weights. If you're building your own model, this one's worth your
    Image
    00:00
    2
  • @datologyai
    DatologyAI
    @datologyai
    21h
    You've probably seen the posts about Thomson, the frontier-competitive legal model @thomsonreuters launched last month with a training run of $450K. DatologyAI was a key partner helping them curate their proprietary data for mid-training. Our case study on the project is out.
    Image
    1
  • @datologyai
    DatologyAI
    @datologyai
    Sep 4
    A 350M model overtrained to 80,000+ tokens per parameter should be a mistake. Nicholas Roberts' ( @nick11roberts ) scaling law says it's optimal. The Train-to-Test (T²) law folds pass@k into pretraining scaling, so once a model reasons at test time, being compute-optimal flips
    Image
    00:00
    2
  • @datologyai
    DatologyAI
    @datologyai
    Aug 26
    Most scaling laws treat training data as a single number: total tokens. So you can't tell which domains help or interfere with each other. In this talk, @kimiahmdh (MIT) shows how to measure the synergy and interference between data domains using a practical method with open
    Image
    00:00
    1
  • @datologyai
    DatologyAI
    @datologyai
    Aug 24
    Congratulations to our partners @thomsonreuters for the launch of Thomson-1.0 today! Leveraging Thomson Reuters' large collection of domain relevant proprietary data curated by DatologyAI, Thomson-1.0-Large was trained for only $450K in compute, can be deployed for a fraction
    @thomsonreuters
    Thomson Reuters
    @thomsonreuters
    Aug 24
    Today at #ILTACON2026, the largest legal technology conference globally, we launched Thomson, our first proprietary large language model, built in-house, and fully owned and controlled by Thomson Reuters. A few lines from today's announcement that stuck with us: “Thomson
    Image
    00:00
    2
Advertisement
Advertisement