Official code for "Towards An End-to-End Framework for Flow-Guided Video Inpainting" (CVPR2022)

Last update: Jan 07, 2023

Overview

E²FGVI (CVPR 2022)

This repository contains the official implementation of the following paper:

Towards An End-to-End Framework for Flow-Guided Video Inpainting
Zhen Li^#, Cheng-Ze Lu^#, Jianhua Qin, Chun-Le Guo^*, Ming-Ming Cheng
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

[Paper] [Demo Video (Youtube)] [演示视频 (B站)] [Project Page (TBD)] [Poster (TBD)]

You can try our colab demo here:

⭐ News

2022.05.15: We release E²FGVI-HQ, which can handle videos with arbitrary resolution. This model could generalize well to much higher resolutions, while it only used 432x240 videos for training. Besides, it performs better than our original model on both PSNR and SSIM metrics. 🔗 Download links: [Google Drive] [Baidu Disk] 🎥 Demo video: [Youtube] [B站]
2022.04.06: Our code is publicly available.

Demo

More examples (click for details):

Coco (click me)	Tennis
Space	Motocross

Overview

🚀 Highlights:

SOTA performance: The proposed E²FGVI achieves significant improvements on all quantitative metrics in comparison with SOTA methods.
Highly effiency: Our method processes 432 × 240 videos at 0.12 seconds per frame on a Titan XP GPU, which is nearly 15× faster than previous flow-based methods. Besides, our method has the lowest FLOPs among all compared SOTA methods.

Work in Progress

Update website page
Hugging Face demo
Efficient inference

Dependencies and Installation

Clone Repo

git clone https://github.com/MCG-NKU/E2FGVI.git

Create Conda Environment and Install Dependencies
```
conda env create -f environment.yml
conda activate e2fgvi
```
- Python >= 3.7
- PyTorch >= 1.5
- CUDA >= 9.2
- mmcv-full (following the pipeline to install)
If the environment.yml file does not work for you, please follow this issue to solve the problem.

Get Started

Prepare pretrained models

Before performing the following steps, please download our pretrained model first.

Model	🔗 Download Links	Support Arbitrary Resolution ?	PSNR / SSIM / VFID (DAVIS)
E²FGVI	[Google Drive] [Baidu Disk]	❌	33.01 / 0.9721 / 0.116
E²FGVI-HQ	[Google Drive] [Baidu Disk]	⭕	33.06 / 0.9722 / 0.117

Then, unzip the file and place the models to release_model directory.

The directory structure will be arranged as:

release_model
   |- E2FGVI-CVPR22.pth
   |- E2FGVI-HQ-CVPR22.pth
   |- i3d_rgb_imagenet.pt (for evaluating VFID metric)
   |- README.md

Quick test

We provide two examples in the examples directory.

Run the following command to enjoy them:

# The first example (using split video frames)
python test.py --model e2fgvi (or e2fgvi_hq) --video examples/tennis --mask examples/tennis_mask  --ckpt release_model/E2FGVI-CVPR22.pth (or release_model/E2FGVI-HQ-CVPR22.pth)
# The second example (using mp4 format video)
python test.py --model e2fgvi (or e2fgvi_hq) --video examples/schoolgirls.mp4 --mask examples/schoolgirls_mask  --ckpt release_model/E2FGVI-CVPR22.pth (or release_model/E2FGVI-HQ-CVPR22.pth)

The inpainting video will be saved in the results directory. Please prepare your own mp4 video (or split frames) and frame-wise masks if you want to test more cases.

Note: E²FGVI always rescales the input video to a fixed resolution (432x240), while E²FGVI-HQ does not change the resolution of the input video. If you want to custom the output resolution, please use the --set_size flag and set the values of --width and --height.

Example:

# Using this command to output a 720p video
python test.py --model e2fgvi_hq --video <video_path> --mask <mask_path>  --ckpt release_model/E2FGVI-HQ-CVPR22.pth --set_size --width 1280 --height 720

Prepare dataset for training and evaluation

Dataset	YouTube-VOS	DAVIS
Details	For training (3,471) and evaluation (508)	For evaluation (50 in 90)
Images	[Official Link] (Download train and test all frames)	[Official Link] (2017, 480p, TrainVal)
Masks	[Google Drive] [Baidu Disk] (For reproducing paper results)

The training and test split files are provided in datasets/<dataset_name>.

For each dataset, you should place JPEGImages to datasets/<dataset_name>.

Then, run sh datasets/zip_dir.sh (Note: please edit the folder path accordingly) for compressing each video in datasets/<dataset_name>/JPEGImages.

Unzip downloaded mask files to datasets.

The datasets directory structure will be arranged as: (Note: please check it carefully)

datasets
   |- davis
      |- JPEGImages
         |- <video_name>.zip
         |- <video_name>.zip
      |- test_masks
         |- <video_name>
            |- 00000.png
            |- 00001.png   
      |- train.json
      |- test.json
   |- youtube-vos
      |- JPEGImages
         |- <video_id>.zip
         |- <video_id>.zip
      |- test_masks
         |- <video_id>
            |- 00000.png
            |- 00001.png
      |- train.json
      |- test.json   
   |- zip_file.sh

Evaluation

Run one of the following commands for evaluation:

 # For evaluating E2FGVI model
 python evaluate.py --model e2fgvi --dataset <dataset_name> --data_root datasets/ --ckpt release_model/E2FGVI-CVPR22.pth
 # For evaluating E2FGVI-HQ model
 python evaluate.py --model e2fgvi_hq --dataset <dataset_name> --data_root datasets/ --ckpt release_model/E2FGVI-HQ-CVPR22.pth

You will get scores as paper reported if you evaluate E²FGVI. The scores of E²FGVI-HQ can be found in [Prepare pretrained models].

The scores will also be saved in the results/<model_name>_<dataset_name> directory.

Please --save_results for further evaluating temporal warping error.

Training

Our training configures are provided in train_e2fgvi.json (for E²FGVI) and train_e2fgvi_hq.json (for E²FGVI-HQ).

Run one of the following commands for training:

 # For training E2FGVI
 python train.py -c configs/train_e2fgvi.json
 # For training E2FGVI-HQ
 python train.py -c configs/train_e2fgvi_hq.json

You could run the same command if you want to resume your training.

The training loss can be monitored by running:

tensorboard --logdir release_model

You could follow this pipeline to evaluate your model.

Results

Quantitative results

Citation

If you find our repo useful for your research, please consider citing our paper:

@inproceedings{liCvpr22vInpainting,
   title={Towards An End-to-End Framework for Flow-Guided Video Inpainting},
   author={Li, Zhen and Lu, Cheng-Ze and Qin, Jianhua and Guo, Chun-Le and Cheng, Ming-Ming},
   booktitle={IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
   year={2022}
}

Contact

If you have any question, please feel free to contact us via zhenli1031ATgmail.com or czlu919AToutlook.com.

License

Licensed under a Creative Commons Attribution-NonCommercial 4.0 International for Non-commercial use only. Any commercial use should get formal permission first.

Acknowledgement

This repository is maintained by Zhen Li and Cheng-Ze Lu.

This code is based on STTN, FuseFormer, Focal-Transformer, and MMEditing.

Official code for "Towards An End-to-End Framework for Flow-Guided Video Inpainting" (CVPR2022)

Related tags

Overview

E2FGVI (CVPR 2022)

⭐ News

Demo

More examples (click for details):

Overview

🚀 Highlights:

Work in Progress

Dependencies and Installation

Get Started

Prepare pretrained models

Quick test

Prepare dataset for training and evaluation

Evaluation

Training

Results

Quantitative results

Citation

Contact

License

Acknowledgement

Owner

Media Computing Group @ Nankai University

[NeurIPS 2021] Garment4D: Garment Reconstruction from Point Cloud Sequences

《DeepViT: Towards Deeper Vision Transformer》(2021)

Open-sourcing the Slates Dataset for recommender systems research

Final project code: Implementing BicycleGAN, for CIS680 FA21 at University of Pennsylvania

No Code AI/ML platform

We have implemented shaDow-GNN as a general and powerful pipeline for graph representation learning. For more details, please find our paper titled Deep Graph Neural Networks with Shallow Subgraph Samplers, available on arXiv (https//arxiv.org/abs/2012.01380).

Unofficial implementation of Pix2SEQ

Code release for DS-NeRF (Depth-supervised Neural Radiance Fields)

Using deep learning to predict gene structures of the coding genes in DNA sequences of Arabidopsis thaliana

The ICS Chat System project for NYU Shanghai Fall 2021

QRec: A Python Framework for quick implementation of recommender systems (TensorFlow Based)

Code for our NeurIPS 2021 paper Mining the Benefits of Two-stage and One-stage HOI Detection

Molecular Sets (MOSES): A Benchmarking Platform for Molecular Generation Models

IJCAI2020 & IJCV 2020 :city_sunrise: Unsupervised Scene Adaptation with Memory Regularization in vivo

Repository for "Improving evidential deep learning via multi-task learning," published in AAAI2022

The code of NeurIPS 2021 paper "Scalable Rule-Based Representation Learning for Interpretable Classification".

HairCLIP: Design Your Hair by Text and Reference Image

Based on Yolo's low-power, ultra-lightweight universal target detection algorithm, the parameter is only 250k, and the speed of the smart phone mobile terminal can reach ~300fps+

SPT_LSA_ViT - Implementation for Visual Transformer for Small-size Datasets

Implementation of Bagging and AdaBoost Algorithm

E²FGVI (CVPR 2022)