Turbo Harness: Instance-Adaptive Harness Optimization

Rutgers University; Red Hat AI Innovation; MIT-IBM Watson AI Lab
Abstract

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

UC Santa Barbara; MIT CSAIL; MIT-IBM Watson AI Lab
Abstract

Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

Google Research; Yale University; Google DeepMind
Abstract

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

EPFL; Apple
Abstract

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Princeton University; NVIDIA; University of Maryland
Abstract

Belief-Aware Multi-Agent Path Finding under Map Uncertainty

Computer Science and Artificial Intelligence Laboratory; Massachusetts Institute of Technology; Symbotic Inc.
Abstract

Learning from Research: Toward Lifelong Agent Harness Evolution

University of California, Santa Barbara; Microsoft
Abstract

Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR

University of Dhaka
Abstract

Agent Error Dataset: Scaling 50,000 Error–Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

Apodex
Abstract

PTNO: Training Neural Operators with Noisy Monte Carlo Estimates for Particle Transport Problems

California Institute of Technology; Yale University; UK Atomic Energy Authority; LIX, CNRS, École polytechnique
Abstract

Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents

The Tesseract Academy
Abstract

What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents

The Hong Kong Polytechnic University
Abstract

AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC

The Hong Kong University of Science and Technology
Abstract

Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking

AIData2Action; TW3 Partners
Abstract

Coverage Before Control: Route-Instruction Grounding and Steering for Controllable Retrosynthesis

The Chinese University of Hong Kong, Shenzhen; Shanghai AI Lab
Abstract

ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible

National University of Singapore; Nanjing University of Science and Technology
Abstract

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Tsinghua University; University of California, Los Angeles; RIKEN AIP; University of Waterloo; Carnegie Mellon University; Yale University; Northwestern University; Zhejiang University; University of California, Berkeley; University of Illinois Chicago; Boston University; The University of Hong Kong; Stanford University; New York University; TokenWave.AI; The University of Tokyo
Abstract

GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning

University of Catania; University of Bologna; ISTC - National Research Council
Abstract

Completion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding

Central South University
Abstract

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

German Research Center for Artificial Intelligence (DFKI); ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center
Abstract

Why ChatPaper

  • Interest-Driven Paper Curation

    Describe your interests as you like, be it keywords or sentence, you’ll get relevant papers every day via AI semantic matching.

  • Top Conferences, One-Stop

    IJCAI, ICML, CVPR, KDD... Access papers from top AI conferences, all in one tool, for free.

  • Paper Management Made Easy

    Keep your research organized with our bookmarking feature. Build a well-structured knowledge hub at your fingertips.

  • Chat With Any Paper

    Got questions? Ask in ChatDOC with one single click. Use AI to pull data, clarify terms, and verify facts with our precise word-level tracing feature.

See how it works

Watch now 1 min