OpenLongTail

Generative Scaling of Long-Tail Driving Data

1Texas A&M University   2NVIDIA   3University of Wisconsin–Madison   4The University of Texas at Austin   5Yale University   6Adobe   7Meta   8University of Delaware   9Stanford University
*Equal contribution    †Corresponding author

Abstract

TL;DR OpenLongTail is an open-scaling generative data engine that converts heterogeneous long-tail driving videos into pose-grounded, multi-view data, enabling scalable VLA policy learning from heterogeneous sources.

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing context. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.

From a Single View to a Full Surround Rig

InputFront
GeneratedCross-Left
GeneratedCross-Right
GeneratedRear-Left
GeneratedRear-Right
GeneratedRear-Tele

OpenLongTail Improves VLA Driving Policies

Closed-Loop Evaluation in AlpaSim

Fine-tuning Alpamayo-R1 with OpenLongTail-synthesized multi-view data improves closed-loop driving robustness on 53 long-tail events (two rollouts each) and approaches training with ground-truth multi-camera capture. AS = AlpaSim Score (higher is better; mean ± 95% bootstrap CI); CR = Collision Rate (lower is better).

Model Additional Training Data Representative Long-tail Events Avg. (N = 106)
10K
Rand
NV-OODWaymo-E2E Complex Int.Cyclists Uncommon Veh.Work Zone
GTSyn GTSyn AS↑CR↓ AS↑CR↓ AS↑CR↓ AS↑CR↓ AS↑CR↓
Alpamayo R10.457 ± 0.02869.6%0.534 ± 0.04050.0%0.575 ± 0.02750.0%0.683 ± 0.02650.0%0.531 ± 0.02658.5%
Alpamayo 1.50.469 ± 0.03073.9%0.347 ± 0.077100.0%0.517 ± 0.03133.3%0.731 ± 0.02250.0%0.501 ± 0.03260.4%
Alpamayo R1 + SFT✓0.659 ± 0.01826.1%0.670 ± 0.0350.0%0.733 ± 0.0170.0%0.936 ± 0.0120.0%0.717 ± 0.02511.3%
✓✓0.743 ± 0.0144.3%0.688 ± 0.0220.0%0.758 ± 0.0140.0%0.938 ± 0.0100.0%0.764 ± 0.0191.9%
✓✓0.747 ± 0.0146.5%0.662 ± 0.0270.0%0.717 ± 0.0142.8%0.935 ± 0.0120.0%0.748 ± 0.0213.8%
✓✓✓0.691 ± 0.01813.0%0.691 ± 0.0250.0%0.735 ± 0.0140.0%0.958 ± 0.0070.0%0.736 ± 0.0245.7%
✓✓✓0.699 ± 0.01713.0%0.689 ± 0.0270.0%0.765 ± 0.0140.0%0.950 ± 0.0090.0%0.749 ± 0.0235.7%

Adding OpenLongTail-synthesized NVIDIA long-tail assets (NV-OOD Syn) raises average AS from 0.717 (nominal-only SFT) to 0.748 and lowers the collision rate from 11.3% to 3.8%, compared with 0.764 AS and 1.9% CR for ground-truth assets (NV-OOD GT). Best and second-best results are shown in bold and underlined; highlighted rows use synthesized data.
SFT: supervised fine-tuning  ·  10K Rand: 10K randomly sampled trajectories  ·  NV-OOD: NVIDIA PAV out-of-distribution subset  ·  GT: ground-truth multi-view  ·  Syn: OpenLongTail-synthesized multi-view.

Generative Scaling Across Diverse Sources

OpenLongTail converts external monocular videos (e.g., Waymo E2E) into target-rig multi-view assets, expanding the long-tail training pool. Combining in-house (NVIDIA PAV) and external synthesized data yields more scene-consistent closed-loop rollouts across uncommon vehicles, cyclists, and complex intersections.

Closed-loop benefits of generative scaling across sources

Closed-loop bird's-eye-view rollouts of Alpamayo-R1 versus variants fine-tuned with OpenLongTail-synthesized assets from NVIDIA PAV data and from PAV combined with an external source (Waymo E2E). Generative scaling expands useful long-tail coverage rather than merely increasing the number of training samples.

Method Overview

OpenLongTail pipeline overview

Given a long-tail front-view driving video, OpenLongTail recovers a metric-scale ego-trajectory and synthesizes five target views under a target camera rig. Cross-left (CL) and cross-right (CR) views are generated with front-anchored view extrapolation, while three rear views are jointly generated with masked view-level attention and road-surface guidance. Together with the input front view, the outputs form synchronized, posed multi-view assets.

Metric Pose Recovery

MapAnything recovers a metric-scale ego-trajectory, stabilized by Kalman + RTS smoothing for jitter-free pose conditioning.

Plücker Ray Geometry

Per-token camera rays encode target-camera geometry so every branch operates in one consistent 3D ray frame.

Temporal Depth Warp

Depth-based reprojection propagates visible front-view evidence into side and rear targets, even with zero overlap.

Cross-View Memory

A directed memory graph lets each target view condition on generated neighbors for seam-level consistency.

BibTeX

@misc{liu2026openlongtailgenerativescalinglongtail,
      title={OpenLongTail: Generative Scaling of Long-Tail Driving Data},
      author={Lulin Liu and Nuo Chen and Yan Wang and Bangya Liu and Wenyan Cong and Hezhen Hu and Boris Ivanovic and Hao Wang and Ziyao Zeng and Xinyu Gong and Yang Zhou and Zixiang Xiong and Dilin Wang and Zhangyang Wang and Weisong Shi and Ruohan Zhang and Marco Pavone and Zhiwen Fan},
      year={2026},
      eprint={2607.09655},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.09655},
}