Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors
Project Page · Paper (PDF) · Supplementary Videos · arXiv coming soon · Code coming soon
InstructVVT is an instruction-driven and reference-guided video virtual try-on framework. Given a source video, a reference garment, and a natural-language instruction, it edits the instructed target while preserving identity, motion, scene structure, and temporal consistency. No masks, poses, parsing maps, DensePose conditions, or garment contours are required at inference time.
This repository currently hosts the project page. The paper PDF and supplementary qualitative videos are available on the page; the arXiv identifier and code release will be added when they are ready.
Dingbao Shao*, Song Wu*, Xinyu Chen, Qian Wang, Jiahang Li, Kuai Jiang, Jiang Lin, Yuhang Liu, Ziyu Chen, Duo Li, Jiaxin Hu, Shengrong Gu, Ziheng Tang, Rongrong Liu, Yanlun Peng, Liang Li, Junlan Feng, Lujia Jin, Ting Zhang, Jian Yang, Zili Yi†
* Equal contribution. † Corresponding author.
The site is dependency-free and served directly from index.html. For local preview:
python3 -m http.server 8000Then open http://localhost:8000.