alphaXiv

Explore

Researchers

Sign In

MCP Server

Autoresearch

Browser Extension

Light themeDark theme

BlogSend Feedback?

Ask questions across all of research

alphaXiv connects papers, researchers, and organizations, grounding the answer in the underlying work.

Alt + Enter to search
Sign up
GLM-5.3: Frontier Coding with Emergent Cyber Capabilities
14 Aug 2026
Z.ai

GLM-5.3 is built on the same base model as GLM-5.2, with every reported gain coming from scaled post-training on long-horizon task environments. It posts open-source SOTA on Terminal Bench 3.0 (28.3 vs 4.6) and Agents' Last Exam (28.5), and a 50% improvement on Z.ai's in-house Code Bench while spending fewer output tokens. Cyber capability grew fastest of all: 84.5% on CyberGym is the best result on that benchmark, and exploitation scores more than doubled. Weights are slated for release two weeks after launch, once safety hardening completes.

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Training AI Scientists to Replicate Research

13 Aug 2026
Damon FalckSamer SabriAnja Surina

Researchers at Inherent developed Faraday, an AI Scientist agent designed to replicate research papers by reproducing experimental figures using a "Coding Agent as a Tool" (CAT) paradigm. Trained on the Replica task space with a novel rubric-based reward system, Faraday demonstrated superior scientific rigor and replication accuracy compared to leading frontier models, including Claude Opus 4.8 and GPT-5.5.

Autoresearch
Paper thumbnail
View PDF
206

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

13 Aug 2026
DreamX TeamRui ChenXiangxiang Chu

DreamX-Phi 1.0 introduces an action-conditioned video world model for robotic manipulation, incorporating geometry-aware SE(3) action representation and comprehensive physical-consistency supervision. This model achieved first place on the WorldArena 2.0 Track 1 leaderboard for video prediction with an EWMScore-P of 60.65 and tied for second on Track 2 for policy training, demonstrating a 67.19% success rate on the "Adjust Bottle" task.

Autoresearch
39
Paper thumbnail
View PDF
249

Researchers to follow

View all
Andrew Ng

Andrew Ng

Managing Partner @ AI Aspire, Managing General Partner @ AI Fund, Executive Chairman @ LandingAI, Founder @ DeepLearning.AI, Adjunct Professor, CS @ Stanford University, Chairman and Co-Founder @ Coursera

Sergey Levine

Sergey Levine

Co-Founder @ Physical Intelligence, Associate Professor, EECS @ UC Berkeley

Yilun Du

Yilun Du

Assistant Professor, CS @ Harvard University, Institute Investigator @ Kempner Institute

Ilya Sutskever

Ilya Sutskever

CEO and Co-Founder @ Safe Superintelligence Inc, Previously Co-Founder and Chief Scientist @ OpenAI

Li Fei-Fei

Li Fei-Fei

Co-Founder and CEO @ World Labs, Founding Co-Director @ Stanford HAI, Sequoia Professor, CS @ Stanford University

Ian Goodfellow

Ian Goodfellow

Co-Founder @ Stealth Startup, Previously Research Scientist @ DeepMind

Chelsea Finn

Chelsea Finn

Co-Founder @ Physical Intelligence, Assistant Professor, CS and EE @ Stanford University

Pieter Abbeel

Pieter Abbeel

Head, Frontier Model Research @ Amazon, Professor, EECS @ UC Berkeley

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

13 Aug 2026
Yaxin LuoHaobin JiangJialv Zou

Meituan and MBZUAI researchers developed AutoDesign, a meta-harness optimization framework enabling design systems to recursively improve their operational components based on human-aligned evaluations. Applied to academic paper-to-poster generation, it achieved a PosterBench Score of 78.32, outperforming other systems by over 7 points, and autonomously generated posters in 40 minutes at under $3 each.

Autoresearch
Paper thumbnail
View PDF
187

V-RAE: Rethinking Video Latent Spaces for Generation

13 Aug 2026
Minghui Guo
Shengqiong WuShengqiong Wu
Hao FeiHao Fei

V-RAE constructs generative latent spaces for video by integrating frozen Vision Foundation Models (VFMs) with a learnable temporal pooling module and a spatiotemporal decoder. This method improves video generation quality, reduces diffusion model training time by up to 6x, and maintains strong semantic information in the latent representations.

Autoresearch
Paper thumbnail
View PDF
175

Intern-S2-Preview: Scientific Agentic Foundation Model

13 Aug 2026
Lei BaiLei Bai
Jiaqi CaoChiyu Chen

The Intern-S2-Preview Team at Shanghai AI Laboratory developed Intern-S2-Preview-397B, a scientific agentic foundation model capable of multimodal scientific understanding, reasoning, and long-horizon tasks through iterative, tool-grounded problem-solving. This model demonstrated competitive or leading performance across diverse scientific, multimodal, and agentic benchmarks, and features specialized modules for time series processing and modular domain adaptation.

Autoresearch
Paper thumbnail
View PDF
156

CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

12 Aug 2026
Zihao YeYingyi HuangHongyi Jin

The CAKE framework, a co-design approach from NVIDIA and Carnegie Mellon University, integrates GPU kernel optimization agents with an evolving compiler harness to overcome limitations of traditional black-box compilation. It enables agents to author high-performance kernels by providing localized diagnostics and a typed, hardware-explicit intermediate representation, achieving significant speedups over baselines and synthesizing frontier kernels for modern GPU architectures.

Autoresearch
Paper thumbnail
View PDF
152

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

12 Aug 2026
Zunhai SuBohan SunXialie Zhuang

This research systematically investigates Massive Activations (MAs) in Hybrid Linear Attention (HLA) Large Language Models, identifying two novel, architecture-aligned morphologies: Pre-attention Spikes (PAS) and Inter-spike Plateaus (ISP). The work demonstrates these patterns' robust recurrence across model scales up to 397B parameters and diverse settings, providing a mechanistic explanation for their emergence and evolution.

Autoresearch
Paper thumbnail
View PDF
193

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

13 Aug 2026
Yuanyang YinGongxuan WangYifan Zhan

Alaya-EVOKE introduces an interactive world model that enables continuous, open-ended virtual world generation by decoupling persistent state from the denoiser and utilizing a long-horizon, dynamically conditioned teacher. The system achieves hour-scale coherent generation with bounded computational costs and supports responsive mid-session interaction and geometric recall.

Autoresearch
Paper thumbnail
View PDF
85

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

13 Aug 2026
Bobo LiHao FeiTianjie Ju

OmniScientist presents an end-to-end AI scientist capable of conducting multidisciplinary research directly from raw, heterogeneous scientific evidence, integrating perception throughout the research lifecycle. It successfully completes full research workflows across diverse scientific domains and modalities, demonstrating improved research quality and multimodal grounding compared to systems relying on pre-processed data.

Autoresearch
Paper thumbnail
View PDF
88

Latent On-Policy Self-Distillation

13 Aug 2026
Guibin ZhangJiayang LyuRan Sun

Latent On-Policy Self-Distillation (LOPD) introduces a framework where the privileged context for self-distillation is a learnable latent representation, moving beyond human-engineered artifacts. This approach consistently enhanced agent performance on tool-use and code generation benchmarks, yielding superior aggregate results and improved sample efficiency across multiple LLM backbones.

Autoresearch
Paper thumbnail
View PDF
98

Researchers to follow

View all
Jingren Zhou

Jingren Zhou

Chief Scientist @ Alibaba Group, Previously Researcher @ Microsoft Research

Song Han

Song Han

Distinguished Scientist, Director of Efficient AI Research @ NVIDIA, Associate Professor, EECS @ Massachusetts Institute of Technology

Linxi "Jim" Fan

Linxi "Jim" Fan

Director & Distinguished Research Scientist @ NVIDIA, Previously CS PhD Student @ Stanford University

Quoc V. Le

Quoc V. Le

Co-Founder @ Discovery Loop, Previously Research Scientist @ Google

Junyang Lin

Junyang Lin

Independent Researcher @ Unaffiliated, Previously Tech Lead @ Alibaba Group

Yejin Choi

Yejin Choi

The Dieter Schwarz Foundation Professor, CS & Senior Fellow, HAI @ Stanford University, Distinguished Scientist, Language and Cognition Research @ NVIDIA

Shuran Song

Shuran Song

Assistant Professor, EE, by courtesy of CS @ Stanford University, Previously Assistant Professor, CS @ Columbia University

Stefano Ermon

Stefano Ermon

CEO & Co-Founder @ Inception, Associate Professor, CS @ Stanford University

Fidelity-Constrained Anchoring for Black-Box Denoisers

13 Aug 2026
Masaki Satoh

Morpho, Inc. developed a fidelity-constrained anchoring framework that post-processes black-box denoiser outputs by linearly blending them with the noisy input. The approach effectively controls output fidelity to the input image, balancing denoising performance and naturalness, with SSIM-based anchoring demonstrating robustness across varying noise levels.

Autoresearch
Paper thumbnail
View PDF
102

StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs

13 Aug 2026
Joya ChenZeyun Zhong
Mike Zheng ShouMike Zheng Shou

STREAMTTT introduces a streaming Video-Language Model that overcomes the perception-memory trade-off in real-time video understanding through a dual-memory architecture and a dedicated real-time QA corpus. Its 4B parameter model achieved a 68.59 two-track average on OVO-Bench, outperforming HERMES-7B by 9.39 points and improving real-time perception by 1.4 points and backward tracing by 3.7 points compared to a matched-scale baseline.

Autoresearch
Paper thumbnail
View PDF
93

Synthetic Persona Pretraining: Alignment from Token Zero

13 Aug 2026
Julian MinderViktor MoskvoretskiiRaghav Singhal

Synthetic Persona Pretraining (SPP) introduces a method to instill desired assistant values and identity into large language models directly from "token zero" during pretraining. This foundational approach yields improved constitution following, enhanced jailbreak robustness, and better alignment to out-of-distribution moral dilemmas, demonstrating a more deeply embedded and robust alignment compared to post-training interventions.

Autoresearch
Paper thumbnail
View PDF
108

Stealing Reasoning Traces from Proprietary LLM APIs

10 Aug 2026
Alexander PanfilovAlexander Panfilov
David SchmotzDavid Schmotz
Ilia ShumailovIlia Shumailov

Researchers uncovered an architectural vulnerability in major proprietary large language model APIs where encrypted internal reasoning traces can be extracted in plaintext by weaker models within the same provider's ecosystem. This enables unauthorized model distillation, secret extraction, jailbreaking, and invisible prompt injection attacks.

Autoresearch
Paper thumbnail
View PDF
7,312

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

13 Aug 2026
Yiwei LiWanli YangHexiang Tan

A systematic evaluation framework was developed for autonomous AI R&D agents, moving beyond single final scores to diagnose performance using process-level metrics, controlled comparisons for experience reuse, and harness impact analysis. The study revealed that current frontier models reliably optimize artifacts but exhibit rare genuine innovation (1.2% novel solutions) and inconsistent performance heavily influenced by bottlenecks in solution framing and feedback control, as well as the effectiveness of experience reuse.

Autoresearch
Paper thumbnail
View PDF
75

Blue Noise as a Lattice Gibbs Ensemble

13 Aug 2026
Zhuoran Yi

Modeling blue noise as a lattice Gibbs ensemble with a generalized Gaussian potential allows for a scalable backward sampling method, which generates high-quality adaptive blue noise for 14K images while maintaining constant memory usage.

Autoresearch
Paper thumbnail
View PDF
47

Full-bandwidth transformer

09 Aug 2026
Xi WangZiyang CaiZheng Zhan

This research introduces the full-bandwidth transformer, which enriches the inter-step feedback channel by fusing the previous top-layer hidden state with the current token embedding. This approach significantly improves data efficiency, enabling a 200B-token model to match or exceed the performance of standard baselines trained with 400B-1T tokens, and enhances inference performance on language model evaluations, math, and coding tasks with minimal overhead.

Autoresearch
Paper thumbnail
View PDF
1,585

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

11 Aug 2026
Kushal ChakrabartiKushal Chakrabarti

Agent instruction files like CLAUDE.md exhibit unbounded growth due to "catastrophic remembering," a process where the original rationale for instructions is lost, making safe deletion difficult. Implementing prompt comments that record outcome-grounded reasoning effectively halts this growth, reducing excess instruction size from 211.3% to 1.4% and improving agent instruction-following correctness by 11.6 percentage points.

Autoresearch
Paper thumbnail
View PDF
953

OCG: Optical Character Grounded Visual Reasoning with Flow Matching

14 Aug 2026
lichen HuangXinrui Wu

The OCG framework from the University of Electronic Science and Technology of China recasts visual reasoning as an image-to-image generation task, representing both questions and answers as optical characters within RGB images. By leveraging a dual-teacher self-distillation scheme, the approach achieved 94.6% exact match accuracy on the CLEVR dataset, closely approaching the 98.1% OCR upper bound.

Paper thumbnail
View PDF
24
There are no more papers matching your filters at the moment.
Sign in

Assistant

Advertisement
Advertisement