Pinned
- GLM 5.3 is the newest member to the cost-accuracy pareto for Terminal-Bench 3.0, replacing Grok 4.6 Impressive improvement from GLM 5.2 (4.6%) to GLM 5.3 (32.4%)!
- Congrats to Z.ai for the strong performance on Terminal-Bench 3.0! One of the biggest pieces of feedback we have gotten for TB3 is to increase the timeouts. We calibrated timeouts against frontier models during development, but inference speed can still be aIntroducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: z.ai/blog/glm-5.3
- run SWE-marathon through harbor!Replying to @rishi_desai2A big criticism of v1.0 was that it was too hard to run. v1.1 fixes that. The full eval now runs with one bash script on vanilla Harbor (no fork!). We also added longer agent timeouts, closed-internet sandboxes, an agent-visible wall-clock timer, and stronger verifiers.




