harbor trial handoff
interview an agent after it completes a benchmark task
- On the average day, agents account for over 70% of @tempo's documentation traffic. Instead of endlessly changing button colors, we built a benchmark to understand how they really see things. Introducing stable-bench-v1 - part of a broad suite of stable coin integration evals
- React 🤝 HarborIntroducing ReactBench A benchmark for coding agents on real React work Models write bad React code - useEffect, slow performance, memory leaks
- New deep research benchmark from our friends at PPLX (built on harbor!)We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer. research.perplexity.ai/articles/wandr…
- Android 🤝 Harbor📊 Your Android Bench July update: 1. Added 8 new models, check out the top of the leaderboard! 2. You can now contribute to the benchmark. 3. We standardized our benchmark by transitioning to the @harborframework. Read about what's new → goo.gle/4p7Mc6G







