alphaXiv

Explore

Researchers

Sign In

MCP Server

Autoresearch

Browser Extension

BlogSend Feedback?

Follow the latest research

alphaXiv connects papers, researchers, and organizations, grounding its answers in the underlying work.

Alt + Enter to search
Sign up

Atria Dawn: The Dawn of Agentic Superintelligence

ImageFDU
Honglin GuoTao GuiJi-Rong WenJi-Rong Wen

The development record shows agents can propose methods and execute revisions, while human researchers retain responsibility for goals, evaluation, and research direction.

14 Sept 2026
405views
Image

FlashREINFORCE: Critic-Free Single-Rollout Asynchronous RL for Agentic Language Models

ImageNVIDIA
Jian HuJian HuYifan ZhangYifan ZhangJan KautzJan Kautz

Critic-free asynchronous training can learn from one rollout per prompt while remaining stable under substantial policy lag and preserving tool use.

14 Sept 2026
6kviews
Image

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

ImageGoogleImageUMD
Tong ZhengXidong WuBenjamin ColemanBenjamin Coleman

Replaying discovery histories lets agents improve exploration strategies cheaply before deploying them, reducing the cost of long-horizon algorithmic and scientific search.

14 Sept 2026
867views4
Image

Researchers to follow

View all
Yann LeCun

Yann LeCun

Executive Chairman

AMI - Advanced Machine Intelligence, Jacob T. Schwartz Professor, CS @ New York University

Kaiming He

Kaiming He

Distinguished Scientist

Google DeepMind, Associate Professor, EECS @ MIT

Ilya Sutskever

Ilya Sutskever

CEO and Co-Founder

Safe Superintelligence Inc

Chelsea Finn

Chelsea Finn

Co-Founder

Physical Intelligence, Assistant Professor, CS and EE @ Stanford University

Alex L. Zhang

Alex L. Zhang

CS PhD Student

Massachusetts Institute of Technology, Research Fellow @ Prime Intellect

John Schulman

John Schulman

Co-Founder and Chief Scientist

Thinking Machines

Andrej Karpathy

Andrej Karpathy

Researcher

Anthropic

Li Fei-Fei

Li Fei-Fei

Co-Founder and CEO

World Labs, Founding Co-Director @ Stanford HAI, Sequoia Professor, CS @ Stanford University

Are you a researcher? Find your profile

Recurrent Looped Transformer

Yifan ZhangYifan Zhang

A recurrent decoder carries computation across every prompt and response token, while exact full-history replay keeps reinforcement-learning policy states aligned with current parameters.

13 Sept 2026
79kviews
Image
Nature Is Our Learning Environment
ImagePeriodic Labs

Periodic Labs describes training an AI system on laboratory data to automate difficult, multiphase XRD interpretation and expand autonomous materials experiments.

15 Sept 2026
539views
Nature Is Our Learning Environment

RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

ImageUC San DiegoImageUIC
Sibo ZhuShicheng FanBiwei HuangBiwei Huang

Agents can adapt to unfamiliar software by exploring, verifying outcomes, and reusing environment-specific memory without changing model parameters.

14 Sept 2026
2kviews
Image

Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning

ImageUPennImageHarvard
Sophia TangSophia TangShiyi Wang

A teacher-free transport map enables language models to generate and iteratively self-correct complete sequences in one or a few evaluations.

15 Sept 2026
616views
Image

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

ImageUniversity of MontrealImageAI2
Michael NoukhovitchMichael NoukhovitchHamish IvisonHamish IvisonAaron CourvilleAaron Courville

Adaptive resampling shifts reinforcement-learning compute from already-solved prompts toward harder ones, improving language models’ ability to solve difficult math and coding tasks.

11 Sept 2026
2kviews
Image

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

ImageGoogle ResearchImageCMU
Honghao LinDavid P. WoodruffVahab MirrokniVahab Mirrokni

A staged multi-agent workflow lets language models explore, decompose, verify, and revise research proofs while preserving objections and partial progress.

15 Sept 2026
278views
Image

Modality-Autoregressive World-Action Models

ImageCMU
Adam HungBardienus P. DuisterhofDeva RamananDeva Ramanan

Predicting structured motion, semantic, and geometric futures before actions improves robot manipulation and lets actionless human videos strengthen learning.

15 Sept 2026
Image

Group Versus Batch Baseline Under a Fixed Sampling Budget

Jian Hu

At equal response budgets, batch or group centering should be chosen from gradient structure and baseline quality, not difficulty distribution shape alone.

15 Sept 2026
136views
Image

Researchers to follow

View all
Ion Stoica

Ion Stoica

Co-Founder & Executive Chairman

Anyscale, Co-Founder & Executive Chairman @ Databricks, Professor, CS @ UC Berkeley

Geoffrey Hinton

Geoffrey Hinton

Emeritus Professor, CS

University of Toronto

Andrew Ng

Andrew Ng

Managing Partner

AI Aspire, Managing General Partner @ AI Fund, Founder @ DeepLearning.AI, Adjunct Professor, CS @ Stanford University, Chairman and Co-Founder @ Coursera

Sergey Levine

Sergey Levine

Co-Founder

Physical Intelligence, Associate Professor, EECS @ UC Berkeley

Christopher D Manning

Christopher D Manning

General Partner

AIX Ventures, Senior Fellow, HAI @ Stanford University

Demis Hassabis

Demis Hassabis

Chair

Google DeepMind, Chief Scientist @ Alphabet, Founder & CEO @ Isomorphic Labs

Yoshua Bengio

Yoshua Bengio

President and Scientific Director

LawZero, Founder and Scientific Advisor @ Mila - Quebec Artificial Intelligence Institute, Canada CIFAR AI Chair @ CIFAR, Full Professor, CS @ Université de Montréal

Yejin Choi

Yejin Choi

The Dieter Schwarz Foundation Professor, CS & Senior Fellow, HAI

Stanford University, Distinguished Scientist, Language and Cognition Research @ NVIDIA

Does Recurrence Pay? A Controlled Evaluation of the Recurrent Looped Transformer

Leon LehmannCasie Nakamura

Controlled small-scale experiments show that per-token recurrence adds substantial training cost without improving language-model quality or long-range prediction.

15 Sept 2026
134views
Image

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

Ling YangLing YangZhenfei YinZhenfei YinYingcheng Wu

The framework enables AI systems to revise research questions, representations, and experiments as external evidence reshapes what remains unknown.

14 Sept 2026
340views
Image

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

ImageUC BerkeleyImageUW
Kaiyuan LiuQiuyang MangLuke ZettlemoyerLuke Zettlemoyer

Across open-ended benchmarks, agents’ marginal gains eventually fall below independent sampling, enabling Elo curves to identify when parallel sessions become more effective.

14 Sept 2026
Image

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

Shuhan XueJianyuan ZhongLing YangLing Yang

Researchers can turn interactive scientific workflows into reusable tasks and rubrics that continually improve an agent’s procedures and problem-solving coverage.

15 Sept 2026
6
Image

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

ImageSJTUImageFDU
Zhenjie YangYideng ZhangHongyang LiHongyang Li

A shared simulation interface lets researchers compare visuo-tactile bimanual manipulation policies across diverse dexterous hands and controlled scene shifts.

14 Sept 2026
Image

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

ImageStanfordImageCMU
Keertana ChidambaramAndrew IlyasAndrew IlyasVasilis SyrgkanisVasilis Syrgkanis

Benign-sounding reasoning injected into context can steer models toward harmful plans that they paraphrase as their own, evading chain-of-thought monitors.

14 Sept 2026
106views
Image

Specifying Reward Functions for RL Without Environment Sampling

ImageStanfordImageUT Austin
Stephane Hatgis-KessellW. Bradley KnoxEmma BrunskillEmma Brunskill

Preference-based reward design can align reinforcement-learning objectives without costly or unsafe environment rollouts, using imagined trajectories and language-model-designed features.

14 Sept 2026
141views
Image

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

ImageMITImageCMU
Xingyang LiDongyun ZouSong HanSong Han

Training-free value smoothing and direct probability encoding make low-bit attention both more faithful and faster for video diffusion models.

14 Sept 2026
Image

MessyMem: Learning-from-Doing Memory for Mobile Manipulation

ImageStanford
Anuva BanwasiWilliam Muckelroy IIIJeannette BohgJeannette Bohg

Robots can reuse interaction outcomes and fine-grained visual memories across rooms and tasks, avoiding repeated exploration during long-horizon manipulation.

15 Sept 2026
Image
There are no more papers matching your filters at the moment.
Sign in

Assistant

Advertisement
Advertisement