Log inSign up
Gert Labs Inc.
38 posts
Gert Labs Inc. profile banner
@GertLabs

Gert Labs Inc.

@GertLabs
Autoscaling unsaturated RL environments and synthetic data. Official website: gertlabs.com
gertlabs.com
Joined April 2026
1
Following
66
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @GertLabs
    Gert Labs Inc.
    @GertLabs
    Aug 20
    We partnered with @kaggle to bring our Adversarial Customer Service evaluation to their massive community of ML practitioners. Adversarial Customer Service is a social intelligence benchmark, where multiple models interact in an environment with their own verifiable goals. The
    Image
    1
  • @GertLabs
    Gert Labs Inc.
    @GertLabs
    19h
    We wanted to see what kinds of LLM jailbreaks exist in the wild, so we built a "autonomous digital asset operations" platform as a honeypot. It was designed as the most tempting target imaginable -- all platform interaction and digital asset management was handled in natural
    Image
  • @GertLabs
    Gert Labs Inc.
    @GertLabs
    Aug 23
    GLM 5.3 results are now live on GBENCH: This is the new frontier (soon to be) open weights model, and by FAR the frontier model in its size class. It's slightly better than Kimi K3 in agentic coding while being faster and cheaper, and it significantly outperforms DeepSeek V4 Pro
    Image
    2
  • @GertLabs
    Gert Labs Inc.
    @GertLabs
    Aug 18
    Qwen 3.8 27B results are now live on GBENCH: It's the new frontier local model, significantly surpassing Qwen 3.6 27B, and comfortably ahead of Glimmer (which is the 2nd best laptop-class model we've tested). It's smarter than GLM 5.1 and the original DeepSeek v4 Flash
    Image
    Image
  • @GertLabs
    Gert Labs Inc.
    @GertLabs
    Aug 10
    Qwen 3.8 Max results are now live: Roughly on par with Kimi K3 in one-shot fluid intelligence, not as strong in agentic coding. Among open weights model, Qwen 3.8 Max and Kimi K3 are in a league of their own, competing with proprietary models like Claude Opus 4.8 and GPT 5.5 on
    Image
Advertisement
Advertisement