Harness-Zero: Harness Distillation via Agent-as-Harness

Peking University; Google; The Hong Kong University of Science and Technology
Abstract

Emergent Collusion in Long-Horizon LLM Agent Interaction

Stanford University; Georgia Tech
Abstract

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Foundation AI, Cisco; Carnegie Mellon University
Abstract

BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction

Yale NLP Lab; New York University; TCS Research
Abstract

A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories

Cornell University; University of Toronto
Abstract

Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models

University of Maryland; Ritual AI; Fudan University; Columbia University
Abstract

Partner-Specific Affective Precision in Social Active Inference

Mission San Jose High School; University of Chicago
Abstract

Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection

Université Côte d'Azur; CNRS; INRIA
Abstract

MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution

University of Oxford; Stanford University
Abstract

GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes

Nanjing University
Abstract

Convex AI Compositionality and the Governance of AI System Populations

University of Zürich; SUPSI; ETH Zürich
Abstract

Construting Reverse Thinking: Developing Large Language Models’ Reverse Thingking Ability

China University of Petroleum (East China); Goertek Inc.; Nankai University
Abstract

Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents

Independent Researcher
Abstract

World State Generator

University of California, Irvine; Northeastern University
Abstract

TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction

Wuhan University; Nanjing Audit University; The University of Manchester; MBZUAI; McGill University; Shanghai Jiao Tong University
Abstract

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Zhejiang University; Alibaba Group
Abstract

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

ITMO University; Hive Trace Lab; Independent Researcher
Abstract

Custom Named Entity Recognition and Topic Classification for Global Health Publications

Faculty of Sciences
Abstract

Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis

Bayer AG
Abstract

The Endless Exam: Mathematical Constructions from Today’s Models toward Superintelligence

The University of Texas at Arlington
Abstract

Why ChatPaper

  • Interest-Driven Paper Curation

    Describe your interests as you like, be it keywords or sentence, you’ll get relevant papers every day via AI semantic matching.

  • Top Conferences, One-Stop

    IJCAI, ICML, CVPR, KDD... Access papers from top AI conferences, all in one tool, for free.

  • Paper Management Made Easy

    Keep your research organized with our bookmarking feature. Build a well-structured knowledge hub at your fingertips.

  • Chat With Any Paper

    Got questions? Ask in ChatDOC with one single click. Use AI to pull data, clarify terms, and verify facts with our precise word-level tracing feature.

See how it works

Watch now 1 min