Follow the latest research
alphaXiv connects papers, researchers, and organizations, grounding its answers in the underlying work.
Researchers to follow
View allAre you a researcher? Find your profile
A shared pixel-space representation lets one model understand, reason about, generate, and edit images while preserving strong multimodal comprehension.
A frozen video world model can explore new camera paths while preserving an event’s appearance, timing, and previously generated scene states.
Multidimensional rarity measures expose culturally specific entities that popularity metrics miss, while reasoning-guided Wikipedia retrieval improves multilingual linking on this neglected long tail.
Researchers to follow
View allRecursive latent-rollout training helps visual world models preserve dynamics-relevant information, enabling long-horizon prediction and robot control across unseen gravitational conditions.
Jointly predicting tokens and multi-token concepts enables language models to reach comparable training loss with substantially fewer tokens than standard architectures.































