Project Page · HuggingFace · Datasets
Bidirectional video models fail to follow simple rules even when trained on enormous amounts of data. We show that serial computation is critical for generating coherent video and use it only in the high-noise regime of video diffusion.
Linux, Python 3.11, and an NVIDIA driver compatible with CUDA 13 are required. Install uv, then install the dependencies from the repository root:
uv sync --lockedVerify the installation:
uv run python -m pytest -q testsGenerate continuations under outputs/demo/. The command downloads the selected
checkpoint and required assets automatically.
uv run python -m tools.demo \
--checkpoint checkpoints/S2PD/dit-b-pixel/conway/model_ema.safetensorsReplace the dataset name in the checkpoint path with any of:
conway, chess, 2048, fifteen_puzzle, tetris, snake, rubiks_3d,
double_pendulum, three_body, colliding_balls, colliding_balls_3d, pong.
Train DiT-B directly on pixels using all visible GPUs. Training logs to W&B by
default. Add --runtime.no-enable-wandb to disable it. See configs/ for other
methods.
bash scripts/train.sh configs/S2PD/DiT-B-pixel.yaml --data.source conwayTraining saves resumable .pth checkpoints and automatically exports
model_ema.safetensors on completion, alongside config.json in the run directory.
Use this file directly for sampling.
Optionally download all released checkpoints (about 84 GiB) and supporting assets:
bash scripts/download_checkpoints.shGenerate validation samples, then evaluate them. The default is 1,024 samples.
Set --sampling.num-samples to change the count. Sampling resumes when rerun
with the same settings. Metrics and diagnostic overlays are saved under
outputs/evaluation/.
bash scripts/sample.sh \
--sampling.checkpoint checkpoints/S2PD/dit-b-pixel/conway/model_ema.safetensors \
--sampling.output-dir outputs/sampling/S2PD/dit-b-pixel/conway
bash scripts/evaluate.sh outputs/sampling/S2PD/dit-b-pixel/conwayTo sample your own training run, replace the checkpoint path with
outputs/training/RUN/model_ema.safetensors and use a new output directory.
To adapt Wan-5B with LoRA fine-tuning, first prepare a video dataset and prepare the video latents using the command below. One measured Rubik’s Cube Real run took 43 minutes and produced 7.7 GiB of latents.
uv run python -m tools.prepare_video_latents --source rubiks_realReplace rubiks_real with any of:
kubric_movi_a, kubric_movi_c, mpm_worlds, rubiks_real.
Obtain MPMWorlds from its original authors.
Then run a demo, train, sample, or evaluate. The demo and sampling commands below use the released checkpoint. CD-FVD requires a GPU and the original validation videos. Matching reference statistics are reused or computed as needed.
# Demo
uv run python -m tools.demo \
--checkpoint checkpoints/S2PD/wan-5b-lora/rubiks_real/model_ema.safetensors
# Training
bash scripts/train.sh configs/S2PD/Wan-5B-LoRA.yaml --data.source rubiks_real
# Sampling
bash scripts/sample.sh \
--sampling.checkpoint checkpoints/S2PD/wan-5b-lora/rubiks_real/model_ema.safetensors \
--sampling.output-dir outputs/sampling/S2PD/wan-5b-lora/rubiks_real
# Evaluation
bash scripts/evaluate_fvd.sh outputs/sampling/S2PD/wan-5b-lora/rubiks_realIf you use S2PD in your research, please cite:
@misc{hu2026s2pd,
title = {{S2PD}: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation},
author = {Hu, Jeffrey and Olmeda Reino, Daniel and Tewari, Ayush},
year = {2026},
eprint = {2610.06847},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2610.06847}
}