[CVPR2021 Oral] End-to-End Video Instance Segmentation with Transformers

Last update: Jan 07, 2023

Related tags

Deep Learning VisTR

Overview

VisTR: End-to-End Video Instance Segmentation with Transformers

This is the official implementation of the VisTR paper:

Installation

We provide instructions how to install dependencies via conda. First, clone the repository locally:

git clone https://github.com/Epiphqny/vistr.git

Then, install PyTorch 1.6 and torchvision 0.7:

conda install pytorch==1.6.0 torchvision==0.7.0

Install pycocotools

conda install cython scipy
pip install -U 'git+https://github.com/cocodataset/cocoapi.git#subdirectory=PythonAPI'
pip install git+https://github.com/youtubevos/cocoapi.git#"egg=pycocotools&subdirectory=PythonAPI"

Compile DCN module(requires GCC>=5.3, cuda>=10.0)

cd models/dcn
python setup.py build_ext --inplace

Preparation

Download and extract 2019 version of YoutubeVIS train and val images with annotations from CodeLab or YoutubeVIS. We expect the directory structure to be the following:

VisTR
├── data
│   ├── train
│   ├── val
│   ├── annotations
│   │   ├── instances_train_sub.json
│   │   ├── instances_val_sub.json
├── models
...

Download the pretrained DETR models on COCO and save it to the pretrained path.

Training

Training of the model requires at least 32g memory GPU, we performed the experiment on 32g V100 card.

To train baseline VisTR on a single node with 8 gpus for 16 epochs, run:

python -m torch.distributed.launch --nproc_per_node=8 --use_env main.py --backbone resnet101/50 --ytvos_path /path/to/ytvos --masks --pretrained_weights /path/to/pretrained_path

Inference

python inference.py --masks --model_path /path/to/model_weights --save_path /path/to/results.json

Models

We provide baseline VisTR models, and plan to include more in future. AP is computed on YouTubeVIS dataset by submitting the result json file to the CodeLab system, and inference time is calculated by pure model inference time (without data-loading and post-processing).

	name	backbone	FPS	mask AP	model	md5
0	VisTR	R50	69.9	34.4	vistr_r50(Please wait)
1	VisTR	R101	57.7	36.5	vistr_r101	2b8d412225121fb1694427ab69a40656

License

VisTR is released under the Apache 2.0 license. Please see the LICENSE file for more information.

Acknowledgement

We would like to thank the DETR open-source project for its awesome work, part of the code are modified from its project.

Citation

Please consider citing our paper in your publications if the project helps your research. BibTeX reference is as follow.

@inproceedings{wang2020end,
  title={End-to-End Video Instance Segmentation with Transformers},
  author={Wang, Yuqing and Xu, Zhaoliang and Wang, Xinlong and Shen, Chunhua and Cheng, Baoshan and Shen, Hao and Xia, Huaxia},
  booktitle =  {Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR)},
  year={2021}
}

[CVPR2021 Oral] End-to-End Video Instance Segmentation with Transformers

Related tags

Overview

VisTR: End-to-End Video Instance Segmentation with Transformers

Installation

Preparation

Training

Inference

Models

License

Acknowledgement

Citation

Owner

Yuqing Wang

Source code for paper "Deep Superpixel-based Network for Blind Image Quality Assessment"

Lbl2Vec learns jointly embedded label, document and word vectors to retrieve documents with predefined topics from an unlabeled document corpus.

Image Matching Evaluation

2.86% and 15.85% on CIFAR-10 and CIFAR-100

This is the repository for Learning to Generate Piano Music With Sustain Pedals

Learn about quantum computing and algorithm on quantum computing

Official implementation of NeurIPS 2021 paper "One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning Objective"

TLoL (Python Module) - League of Legends Deep Learning AI (Research and Development)

Implementation of QuickDraw - an online game developed by Google, combined with AirGesture - a simple gesture recognition application

Resources for the "Evaluating the Factual Consistency of Abstractive Text Summarization" paper

From Canonical Correlation Analysis to Self-supervised Graph Neural Networks

We will release the code of "ConTNet: Why not use convolution and transformer at the same time?" in this repo

A simple command line tool for text to image generation, using OpenAI's CLIP and a BigGAN.

An example of semantic segmentation using tensorflow in eager execution.

All public open-source implementations of convnets benchmarks

WPPNets: Unsupervised CNN Training with Wasserstein Patch Priors for Image Superresolution

Personalized Transfer of User Preferences for Cross-domain Recommendation (PTUPCDR)

This is the implementation of "SELF SUPERVISED REPRESENTATION LEARNING WITH DEEP CLUSTERING FOR ACOUSTIC UNIT DISCOVERY FROM RAW SPEECH" submitted to ICASSP 2022

Microsoft Cognitive Toolkit (CNTK), an open source deep-learning toolkit

This repo is a C++ version of yolov5_deepsort_tensorrt. Packing all C++ programs into .so files, using Python script to call C++ programs further.