1. X
  2. Appen Research
Log inSign up
Appen Research
37 posts
Image
user avatar
Appen Research
@AppenResearch
Human data for frontier AI. Research and insights from Appen.
Joined May 2026
188
Following
241
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Appen Research
    @AppenResearch
    Jun 25
    New benchmark: RADAR evaluates 21 frontier models on vulnerability identification across 41 real-world XBOW codebases, scored against human expert ground truth. Primary metric is recall (asymmetric cost structure in security). Top result: 62.4%. The recall/F1 ranking inversion
    Image
    610
  • user avatar
    Appen Research
    @AppenResearch
    19h
    An agent can pass a benchmark and still fail the test. For Laguna S 2.1, @poolsideai evaluated full trajectories, not just final scores. @AppenResearch worked with Poolside on aspects of the reward hacking detection used during the evaluation process.
    Image
    2.8K
  • user avatar
    Appen Research
    @AppenResearch
    Jul 27
    🎧 Now on Spotify. What happens when AI agents stop worrying about context limits and start solving problems? On the latest episode of The Data Layer, @alex_whedon, Co-founder & CTO of @subquadratic, and @SergioBrucc, VP of Delivery at @AppenResearch, to discuss: • Linear vs.
    399
  • user avatar
    Appen Research
    @AppenResearch
    Jul 23
    Great to see the @poolsideai team release Laguna S 2.1. A fully open 118B MoE model with 8B active parameters per token, up to a 1M-token context window, and designed for long-horizon agentic coding. Looking forward to seeing what the community builds with it. #AI #AgenticAI
    user avatar
    Poolside
    @poolsideai
    Jul 21
    Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many
    Image
    00:00
    284
  • user avatar
    Appen Research
    @AppenResearch
    Jul 23
    Listen to The Data Layer by Appen on @Spotify Is Your Speech AI Actually Good? | Appen × Hugging Face | Benchmarked | The Data Layer Ep. 1 @huggingface @EricBezzam @SergioBrucc
    97
  • See @AppenResearch's full profile

    Sign up
    Log in
Advertisement
Advertisement