1. X
  2. Appen Research
Log inSign up
Appen Research
75 posts
Appen Research profile banner
user avatar

Appen Research

@AppenResearch
Human data for frontier AI. Research and insights from Appen.
Joined May 2026
211
Following
282
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    Appen Research
    @AppenResearch
    5h
    Pass/Fail isn’t enough for Physical AI. A robot can complete a household task while taking an inefficient trajectory, making unstable grasps, or nearly dropping the object. Those behaviors matter and binary benchmarks don’t capture them. See how Appen partnered with a frontier
    Image
    GIF
  • user avatar
    Appen Research
    @AppenResearch
    Aug 27
    VLM judges can be right about the details and still get the overall judgment wrong. If the important moment never gets retrieved, better reasoning won’t fix it. And even when a judge catches a real omission, the bigger question is whether it actually mattered. New from
    Image
  • user avatar
    Appen Research
    @AppenResearch
    Aug 26
    Scientific reasoning isn’t a data volume problem. It demands expert knowledge density. That’s why the Appen Research team built 35+ Academic STEM Off-the-Shelf Datasets across adjacent expert domains, expert-authored corpora sourced from peer-reviewed journals and professional
    Image
    Image
    Image
  • user avatar
    Appen Research
    @AppenResearch
    Aug 26
    This is exactly what we’re seeing in practice. Expert analysis of where and why agents fail can become the feedback that improves how they reason the next time around. Really interesting paper from Google showing this loop in action.
    user avatar
    marfin
    @marfinxx
    Aug 25
    This Google AI research paper is f*cking brilliant A new research paper proves that distilling high-level strategies from failed and successful trajectories transforms static LLMs into self-improving agents ReasoningBank memory extraction, memory-aware test-time scaling
    Image
    Image
  • user avatar
    Appen Research
    @AppenResearch
    Aug 25
    An LLM judge is only as reliable as the language it’s judging in. The MMLU-ProX benchmark, which evaluated 36 frontier models across 29 languages, found performance gaps of up to 24.3% between high and low-resource languages. Appen's Multilingual LLM-as-a-Judge Managed Service
    Image
Advertisement
Advertisement