1. X
  2. Archiki Prasad
Log inSign up
Archiki Prasad
627 posts
Image
user avatar
Archiki Prasad
@ArchikiPrasad
Research Scientist @GoogleDeepMind | Ph.D. from @unccs | @Apple AI/ML Scholar | Prev (intern): GDM, @AIatMeta (FAIR), @allenai_org
New York City, NY
archiki.github.io
Joined December 2016
1,062
Following
2,266
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Archiki Prasad
    @ArchikiPrasad
    May 7
    🎉 Excited to share that our work on intrinsic dimensionality of reasoning has been accepted to #ICML2026 as a ✨spotlight✨ (top 2.2%)! We analyze the effectiveness of teaching a model how to reason via the lens of intrinsic dimensionality (the minimum effective capacity a
    user avatar
    Archiki Prasad
    @ArchikiPrasad
    Feb 11
    🚨Excited to share our new work viewing reasoning strategies as teaching tools: for fixed target model, which CoT strategies best support learning and generalization? ✨Our answer is intrinsic dimensionality (minimum effective capacity a model needs to solve the task). Somewhat
    Image
    21K
  • user avatar
    Archiki Prasad
    @ArchikiPrasad
    May 21
    What makes a good token-level learning signal for LLM reasoning? In ✨AVSD ✨, we condition on multiple types of privileged information (e.g., ground truth answer, partial solutions) as dense teachers: leveraging what different teacher views agree on, while retaining useful
    user avatar
    Duy Nguyen
    @duynguyen772
    May 21
    Sparse binary rewards bottleneck LLM RL, motivating the use of privileged information in self-distillation as dense teachers. How can we use and balance multiple types of privileged info: leveraging stable cross-view info, while preserving view-specific info? Current on-policy
    Image
    4K
  • user avatar
    Archiki Prasad
    @ArchikiPrasad
    May 13
    🚨Excited to share ✨Agent-BRACE✨, our new work on belief state modeling for LLM agents in long-horizon tasks! 🔸Most agents either suffer from growing context window or compress history into summaries, however, they do not explicitly track what the agent does not know.
    user avatar
    Joykirat
    @joykiratsingh
    May 13
    🚨Excited to announce Agent-BRACE! LLM agents in long-horizon POMDPs either blow up their context with raw history or summarize it, discarding uncertainty by collapsing belief into a point estimate. Agent-BRACE decouples the agent into belief state + policy models, jointly
    Image
    7.4K
  • user avatar
    Archiki Prasad
    @ArchikiPrasad
    Apr 10
    🎉 Excited to share that PRInTS has been accepted to #ACL2026 (Main Conference)! PRInTS is a generative PRM that improves agents on long-horizon info-seeking tasks, yielding +9.3% (absolute) gain in avg. accuracy across GAIA, FRAMES & WebWalker! More details in the thread 🧵⬇️
    user avatar
    Archiki Prasad
    @ArchikiPrasad
    Nov 25, 2025
    🚨 Excited to announce ✨PRInTS✨, a generative Process Reward Model (PRM) that improves agent’s long-horizon info-seeking via info-gain scoring + summarization. PRInTS guides open + specialized agents with major boosts 👉+9.3% avg. w/ Qwen3-32B across GAIA, FRAMES &
    Image
    7.3K
  • user avatar
    Archiki Prasad
    @ArchikiPrasad
    Apr 7
    📢 Excited to share our new work on Cog-DRIFT! The core idea is quite simple yet effective: when a problem is too hard for a model to learn from directly, reformat it into an easier variant (MCQ, cloze) first! Key highlights: 🔸Task reformulation unlocks learning from
    user avatar
    Justin Chih-Yao Chen
    @cyjustinchen
    Apr 7
    🚨Cog-DRIFT: Breaking the Exploration Barrier in RLVR RLVR has pushed LLM reasoning forward BUT hits a ceiling: if a model can't solve a problem (rollouts never succeed), it gets 0 learning signal 👉Hard problems stay unsolved, and training stalls. We introduce✨Cog-DRIFT✨to
    Image
    1.9K

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement