Harness-Zero: Harness Distillation via Agent-as-Harness
Emergent Collusion in Long-Horizon LLM Agent Interaction
Et Tu, Brute? Economic Misalignment in Personal AI Agents
BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction
A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories
Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models
Partner-Specific Affective Precision in Social Active Inference
Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection
MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution
GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes
Convex AI Compositionality and the Governance of AI System Populations
Construting Reverse Thinking: Developing Large Language Models’ Reverse Thingking Ability
Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents
ワールドステート生成器
World State Generator
TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction
Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents
DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
Custom Named Entity Recognition and Topic Classification for Global Health Publications
Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis
The Endless Exam: Mathematical Constructions from Today’s Models toward Superintelligence
