Project page · Paper (coming soon) · Data (coming soon)
We study the problem of generating parallel worlds of a flight — visually distinct videos that share the same flight trajectory — from pre-trained video generative models. Controlling such generation presents an inherent trade-off: camera parameters specify motion precisely but are foreign to video models, while visual references are easy to condition on but reveal the original world. WorldSlider is a general generative pipeline for synthesizing diverse visual worlds along prescribed flight trajectories. Central to it is FlightGrid, an appearance-free visual control signal constructed automatically from metric camera poses, which encodes both the trajectory and its time-varying free-space corridor in a form native to video generative models. To make generation efficient and further improve its fidelity to the prescribed trajectory, we augment few-step distillation with reward learning from a geometry-based camera-motion verifier.
Code release is in preparation. This repository is currently a placeholder — nothing has been uploaded yet. In the meantime the project page carries the qualitative results, the method figures, the ablations and the camera-control comparison.
| FlightGrid | Renderer for the appearance-free control signal, built from metric 6 DoF poses |
| Control adaptation | ControlNet-style control branch over a frozen pre-trained video DiT |
| Verifier-guided DMD | Few-step distillation reweighted by a source-calibrated geometry verifier |
| Evaluation | The camera-control benchmark and the scoring pipeline behind Table 1 |
| Data | Curated trajectory–video pairs with metric poses |
@inproceedings{worldslider,
title = {Generating Parallel Worlds of a Flight},
author = {Yan, Keyu and Cao, Wenhan and Wang, Shenao and Wu, Junke and Chen, Xuyang and Zhao, Lin},
booktitle = {TODO},
year = {2027}
}To be decided before the code release.
