1. X
  2. Jacky Kwok
Log inSign up
Jacky Kwok
94 posts
user avatar
Jacky Kwok
@jackyk02
Stanford CS PhD | Berkeley EECS
Palo Alto, CA
linkedin.com/in/jackykwok02/
Joined June 2025
906
Following
700
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Jacky Kwok
    @jackyk02
    Jul 8
    How can we extract richer signals from AI Feedback? Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀 The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take
    Image
  • user avatar
    Jacky Kwok
    @jackyk02
    Jul 21
    Thanks to @whoisnnamdi and @lightspeedvp for the shout-out to LLM-as-a-Verifier on the Lightwork podcast! Check out the episode for a great explanation of "loop engineering" and agent verification:
    Image
    The $20,000 Robot, AI's Essay Wars & Loop Engineering Explained | Lightwork
    From youtube.com
  • user avatar
    Jacky Kwok
    @jackyk02
    Jun 3
    Excited to share that CoVer-VLA has been selected as the Best Paper Finalist at the CVPR 2026 Scalable Robot Learning Workshop 🤖 I’ll be talking about how verification can be scaled 🚀 for robots—both during training and test-time! 📍 Denver Convention Center, Room 610 🕔 June
    Image
    00:00
    Image
    user avatar
    Jacky Kwok
    @jackyk02
    Feb 24
    Introducing CoVer-VLA💫— a contrastive verifier + hierarchical test-time scaling framework for VLAs! - Lightweight 1B verifier 🧠 - Outperforms π₀ & π₀.₅ 🦾 - Trained on Bridge & DROID 🤖 Turns out scaling verification > scaling policy learning for VLA alignment! 🧵👇 🌐
  • user avatar
    Jacky Kwok
    @jackyk02
    Apr 9
    We release LLM-as-a-Verifier 🧠: A general-purpose verification framework that achieves SOTA 👑 on Terminal-Bench 2 (86.4%) and SWE-Bench Verified (77.8%) by scaling: - scoring granularity - repeated verification - criteria decomposition 📄 Blog & Code: llm-as-a-verifier.notion.site
    Image
  • user avatar
    Jacky Kwok
    @jackyk02
    Feb 24
    Introducing CoVer-VLA💫— a contrastive verifier + hierarchical test-time scaling framework for VLAs! - Lightweight 1B verifier 🧠 - Outperforms π₀ & π₀.₅ 🦾 - Trained on Bridge & DROID 🤖 Turns out scaling verification > scaling policy learning for VLA alignment! 🧵👇 🌐
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement