Log inSign up
Zaid Khan
703 posts
@codezakh

Zaid Khan

@codezakh
NDSEG Fellow / PhD @uncnlp with @mohitban47 working on automating env/data generation + program synthesis formerly @allenai @neclabsamerica
Boston, USA
zaidkhan.me
Joined June 2023
1,390
Following
697
Followers
1
Subscription
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @codezakh
    Zaid Khan
    @codezakh
    Jun 2
    Can an LLM act as a selective model of a GPU during evolutionary search, by reasoning + forecasting a kernel’s runtime but deferring to a GPU when unsure? We produced 12k kernels + runtimes from evolutionary search, costing 400M reasoning tokens + 600 GPU-hours to answer this.
    Image
    5
  • @codezakh
    Zaid Khan
    @codezakh
    Jun 10
    In our latest work, we train an open-weights critic that monitors + guides frontier model agents on long-horizon GUI tasks. Key ideas: track visual UI changes caused by agent (don't trust the agent's intents) + monitor macro actions / goals (don't let it go down rabbitholes).
    @hyunji_amy_lee
    hyunji amy lee
    @hyunji_amy_lee
    Jun 10
    🚨 Introducing HiViG, a test-time intervention framework for long-horizon GUI tasks. By tracking history & verifying actions w/ visual grounding, HiViG boosts performance across diverse GUI environments even for strong policies where existing critics often degrade performance.
    Image
  • @codezakh
    Zaid Khan
    @codezakh
    Jun 2
    Appreciate the shoutout @_akhaliq for our work on "GPU Forecasters" exploring whether language models can act as selective surrogates for GPU kernel optimization! Details in our thread: x.com/codezakh/statu…
    @_akhaliq
    AK
    @_akhaliq
    Jun 2
    GPU Forecasters Language Models as Selective Surrogates for Kernel Runtime Optimization
    Image
  • @codezakh
    Zaid Khan
    @codezakh
    May 21
    We’ve been working on a way to get better on-policy token-level rewards for LLMs + RL! Self-distillation gives token-level rewards, using divergence against a teacher policy given privileged info (i.e true final answer). What if you could use multiple forms of privileged info?
    @duynguyen772
    Duy Nguyen
    @duynguyen772
    May 21
    Sparse binary rewards bottleneck LLM RL, motivating the use of privileged information in self-distillation as dense teachers. How can we use and balance multiple types of privileged info: leveraging stable cross-view info, while preserving view-specific info? Current on-policy
    Image
    1
  • @codezakh
    Zaid Khan
    @codezakh
    May 20
    Continually updated envs (e.g. Git repo histories, evolving docs) are central to knowledge work. Reasoning about these requires long context understanding + resolving temporally distributed / interfering changes to the env state. How well do LLM agents / memory systems do? 🧵👇
    @hyunji_amy_lee
    hyunji amy lee
    @hyunji_amy_lee
    May 20
    LLM agents & memory systems operate in continuously updated environments (Git repos, evolving docs). They must process long contexts, recover earlier information, and reason over many updates that create interference between old and new information. How well do they handle this?
    Image
Advertisement
Advertisement