1. X
  2. Daniel Fried
Log inSign up
Daniel Fried
973 posts
Image
user avatar
Daniel Fried
@dan_fried
Assistant prof. @LTIatCMU @SCSatCMU. Working on NLP: LLM agents, language-to-code, applied pragmatics, grounding.
Pittsburgh, PA
dpfried.github.io
Joined August 2013
917
Following
4,396
Followers
RepliesRepliesMediaMedia
  • user avatar
    Daniel Fried
    @dan_fried
    Jul 2
    We're creating a new course on AI Agents at CMU this Fall! We’re aiming to give students hands-on experience: from building agentic harnesses and evals to training with RL. Check out our course site for the full schedule: cmu-agents.com
    user avatar
    Graham Neubig
    @gneubig
    Jul 2
    This Fall at CMU we're teaching a new course on AI Agents! The goal is that you learn how to create a scaffold, build evals, and train an agentic LLM using RL. We'll try to balance theory and practice, and introduce modern frameworks and best practices.
    Image
    50K
  • user avatar
    Daniel Fried
    @dan_fried
    Jun 3
    New work: a simple and general multi-agent computer use framework. It uses a manager to plan and re-plan by creating a task DAG, with subagents for parallel execution. It improves success rate across benchmarks, and substantially improves efficiency on long-horizon tasks.
    Image
    00:00
    Image
    00:23
    user avatar
    Jing Yu Koh
    @kohjingyu
    Jun 3
    Computer use agents are slow and brittle. The fix isn’t just stronger models, but also deploying them as multi-agent systems. MACU is a general Multi-Agent Computer Use framework that consistently lifts success rates by 3.4-25.5% and is up to 1.5x faster on long-horizon tasks.🧵
    4.9K
  • user avatar
    Daniel Fried
    @dan_fried
    Apr 29
    How successfully -- and efficiently! -- can agents carry out long-horizon tasks on the web? We built a benchmark of ~200 multi-site tasks, based on people's real browsing history. Many of them take hours to solve. Paper: odysseys-website.pages.dev Led by @JangLawrenceK and
    user avatar
    Jing Yu Koh
    @kohjingyu
    Apr 29
    One of the things I’m most excited about this year is building agents that can work productively for hours, days, or weeks. Coding agents are starting to become very competent at this, but what about computer use agents? Our new benchmark, Odysseys (co-led with @JangLawrenceK)
    Image
    00:00
    14K
  • user avatar
    Daniel Fried
    @dan_fried
    Apr 24
    Also at #ICLR2026: a new benchmark for coding agents that implement and run experiments from papers. Masking regions of code gives us a knob to control difficulty of the task (still verifiable!) Paper: arxiv.org/abs/2506.19724 Work with @j1mk1m1016, Alex Wilf, and @lpmorency
    user avatar
    James Kim
    @j1mk1m1016
    Apr 24
    🚀 Excited to share our ICLR 2026 paper: "From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking"! Work with Alex Wilf, LP Morency, @dan_fried Check out the project here! iclr.cc/virtual/2026/p…
    Image
    4.5K
  • user avatar
    Daniel Fried
    @dan_fried
    Apr 24
    This morning (Fri) at #ICLR2026, check out Andy's work on ConflictScope: determining how an LLM prioritizes between a set of user-provided values, by generating scenarios where the values are in conflict. P4-#4105
    user avatar
    Andy Liu
    @uilydna
    Apr 20
    I'll be in Rio this week for #ICLR2026 to present "Generative Value Conflicts Reveal LLM Priorities" (Friday morning, P4-#4105). Happy to chat anything related to LLM alignment, human-AI interaction, or multi-agent systems - feel free to DM if interested!
    Image
    2.1K

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement