1. X
  2. Appen Research
Log inSign up
Appen Research
38 posts
Image
user avatar
Appen Research
@AppenResearch
Human data for frontier AI. Research and insights from Appen.
Joined May 2026
208
Following
246
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Appen Research
    @AppenResearch
    Jun 25
    New benchmark: RADAR evaluates 21 frontier models on vulnerability identification across 41 real-world XBOW codebases, scored against human expert ground truth. Primary metric is recall (asymmetric cost structure in security). Top result: 62.4%. The recall/F1 ranking inversion
    Image
  • user avatar
    Appen Research
    @AppenResearch
    Aug 3
    "Bring 100 experts into the room, and you'll get 100 different opinions." At SlatorCon London 2026, Appen's @SergioBrucc discussed why domains like coding, STEM, healthcare, legal, and HR require expert judgment, and why expertise must span cultures, languages, and modalities.
    Image
  • user avatar
    Appen Research
    @AppenResearch
    Jul 30
    An agent can pass a benchmark and still fail the test. For Laguna S 2.1, @poolsideai evaluated full trajectories, not just final scores. @AppenResearch worked with Poolside on aspects of the reward hacking detection used during the evaluation process.
    Image
  • user avatar
    Appen Research
    @AppenResearch
    Jul 27
    🎧 Now on Spotify. What happens when AI agents stop worrying about context limits and start solving problems? On the latest episode of The Data Layer, @alex_whedon, Co-founder & CTO of @subquadratic, and @SergioBrucc, VP of Delivery at @AppenResearch, to discuss: • Linear vs.
  • user avatar
    Appen Research
    @AppenResearch
    Jul 23
    Great to see the @poolsideai team release Laguna S 2.1. A fully open 118B MoE model with 8B active parameters per token, up to a 1M-token context window, and designed for long-horizon agentic coding. Looking forward to seeing what the community builds with it. #AI #AgenticAI
    user avatar
    Poolside
    @poolsideai
    Jul 21
    Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many
    Image
    00:00

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement