Code for Referring Image Segmentation via Cross-Modal Progressive Comprehension, CVPR2020.

Last update: Dec 01, 2022

Related tags

Deep Learning CMPC-Refseg

Overview

CMPC-Refseg

Code of our CVPR 2020 paper Referring Image Segmentation via Cross-Modal Progressive Comprehension.

Shaofei Huang*, Tianrui Hui*, Si Liu, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, Bo Li (* Equal contribution)

Interpretation of CMPC.

(a) Input referring expression and image.
(b) The model first perceives all the entities described in the expression based on entity words and attribute words, e.g., “man” and “white frisbee” (orange masks and blue outline).
(c) After finding out all the candidate entities that may match with input expression, relational word “holding” can be further exploited to highlight the entity involved with the relationship (green arrow) and suppress the others which are not involved.
(d) Benefiting from the relation-aware reasoning process, the referred entity is found as the final prediction (purple mask).

Experimental Results

We modify the way of feature concatenation in the end of CMPC module and achieve higher performances than the results reported in our paper. New experimental results are summarized in the table bellow. You can download our trained checkpoints to test on the four datasets. The link to the checkpoints is: Baidu Drive, pswd: jjsf.

Method	UNC val	UNC testA	UNC testB	UNC+ val	UNC+ testA	UNC+ testB	G-Ref val	ReferIt test
STEP-ICCV19 [1]	60.04	63.46	57.97	48.19	52.33	40.41	46.40	64.13
Ours-CVPR20	61.36	64.53	59.64	49.56	53.44	43.23	49.05	65.53
Ours-Updated	62.47	65.08	60.82	50.25	54.04	43.47	49.89	65.58

Setup

We recommended the following dependencies.

Python 2.7
TensorFlow 1.5
Numpy
pydensecrf

This code is derived from RRN [2]. Please refer to it for more details of setup.

Data Preparation

Dataset Preprocessing

We conduct experiments on 4 datasets of referring image segmentation, including UNC, UNC+, Gref and ReferIt. After downloading these datasets, you can run the following commands for data preparation:

python build_batches.py -d Gref -t train
python build_batches.py -d Gref -t val
python build_batches.py -d unc -t train
python build_batches.py -d unc -t val
python build_batches.py -d unc -t testA
python build_batches.py -d unc -t testB
python build_batches.py -d unc+ -t train
python build_batches.py -d unc+ -t val
python build_batches.py -d unc+ -t testA
python build_batches.py -d unc+ -t testB
python build_batches.py -d referit -t trainval
python build_batches.py -d referit -t test

Glove Embedding

Download Gref_emb.npy and referit_emb.npy and put them in data/. We provide download link for Glove Embedding here: Baidu Drive, password: 2m28.

Training

Train on UNC training set with:

python -u trainval_model.py -m train -d unc -t train -n CMPC_model -emb -f ckpts/unc/cmpc_model

Testing

Test on UNC validation set with:

python -u trainval_model.py -m test -d unc -t val -n CMPC_model -i 700000 -c -emb -f ckpts/unc/cmpc_model

CMPC for video referring segmentation

We release video version code for CMPC on A2D dataset under CMPC_video/.

Reference

[1] Chen, Ding-Jie, et al. "See-through-text grouping for referring image segmentation." Proceedings of the IEEE International Conference on Computer Vision. 2019.

[2] Li, Ruiyu, et al. "Referring image segmentation via recurrent refinement networks." Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018.

Citation

If our CMPC is useful to your research, please consider citing:

@inproceedings{huang2020referring,
  title={Referring Image Segmentation via Cross-Modal Progressive Comprehension},
  author={Huang, Shaofei and Hui, Tianrui and Liu, Si and Li, Guanbin and Wei, Yunchao and Han, Jizhong and Liu, Luoqi and Li, Bo},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages={10488--10497},
  year={2020}
}

Code for Referring Image Segmentation via Cross-Modal Progressive Comprehension, CVPR2020.

Related tags

Overview

CMPC-Refseg

Interpretation of CMPC.

Experimental Results

Setup

Data Preparation

Training

Testing

CMPC for video referring segmentation

Reference

Citation

Owner

spyflying

Codecov coverage standard for Python

A method to perform unsupervised cross-region adaptation of crop classifiers trained with satellite image time series.

Official Implementation of CoSMo: Content-Style Modulation for Image Retrieval with Text Feedback

The codebase for Data-driven general-purpose voice activity detection.

Pytorch implementation of Nueral Style transfer

Code for "On the Effects of Batch and Weight Normalization in Generative Adversarial Networks"

Pytorch implementation of the paper "COAD: Contrastive Pre-training with Adversarial Fine-tuning for Zero-shot Expert Linking."

Predicting future trajectories of people in cameras of novel scenarios and views.

Supervised Contrastive Learning for Product Matching

Implementation of H-Transformer-1D, Hierarchical Attention for Sequence Learning

4D Human Body Capture from Egocentric Video via 3D Scene Grounding

Starter kit for getting started in the Music Demixing Challenge.

improvement of CLIP features over the traditional resnet features on the visual question answering, image captioning, navigation and visual entailment tasks.

Minimisation of a negative log likelihood fit to extract the lifetime of the D^0 meson (MNLL2ELDM)

GLODISMO: Gradient-Based Learning of Discrete Structured Measurement Operators for Signal Recovery

library for nonlinear optimization, wrapping many algorithms for global and local, constrained or unconstrained, optimization

BEAS: Blockchain Enabled Asynchronous & Secure Federated Machine Learning

Airborne Optical Sectioning (AOS) is a wide synthetic-aperture imaging technique

Adaout is a practical and flexible regularization method with high generalization and interpretability

Pomodoro timer that acknowledges the inexorable, infinite passage of time