Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Last update: Nov 25, 2022

Related tags

Deep Learning SGN

Overview

Semantic Grouping Network for Video Captioning

Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang D. Yoo. AAAI 2021. [arxiv]

Environment

Ubuntu 16.04
CUDA 9.2
cuDNN 7.4.2
Java 8
Python 2.7.12
- PyTorch 1.1.0
- Other python packages specified in requirements.txt

Usage

1. Setup

$ pip install -r requirements.txt

2. Prepare Data

Download the GloVe Embedding from here and locate it at data/Embeddings/GloVe/GloVe_300.json.
Extract features from datasets and locate them at data/ /features/ .hdf5.

e.g. ResNet101 features of the MSVD dataset will be located at data/MSVD/features/ResNet101.hdf5.

I refer to this repo for extracting the ResNet101 features, and this repo for extracting the 3D-ResNext101 features.
Split the features into train, val, and test sets by running following commands.
```
$ python -m split.MSVD
$ python -m split.MSR-VTT
```

You can skip step 2-3 and download below files

MSVD
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]
MSR-VTT
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]

3. Prepare The Code for Evaluation

Clone the evaluation code from the official coco-evaluation repo.

$ git clone https://github.com/tylin/coco-caption.git
$ mv coco-caption/pycocoevalcap .
$ rm -rf coco-caption

4. Extract Negative Videos

$ python extract_negative_videos.py

or you can skip this step as the output files are already uploaded at data/ /metadata/neg_vids_ .json

5. Train

$ python train.py

You can change some hyperparameters by modifying config.py.

Pretrained Models - SGN(R101+RN)

*Disclaimer: The models above do not have the same weight as the models used in the paper (I trained them again because I lost).

6. Evaluate

$ python evaluate.py --ckpt_fpath

License

The source-code in this repository is released under MIT License.

Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Related tags

Overview

Semantic Grouping Network for Video Captioning

Environment

Usage

1. Setup

2. Prepare Data

3. Prepare The Code for Evaluation

4. Extract Negative Videos

5. Train

6. Evaluate

License

Owner

Hobin Ryu

ZeroVL - The official implementation of ZeroVL

FCAF3D: Fully Convolutional Anchor-Free 3D Object Detection

Toward Multimodal Image-to-Image Translation

A robust camera and Lidar fusion based velocity estimator to undistort the pointcloud.

SparseInst: Sparse Instance Activation for Real-Time Instance Segmentation, CVPR 2022

The Unreasonable Effectiveness of Random Pruning: Return of the Most Naive Baseline for Sparse Training

Implementation of Multistream Transformers in Pytorch

[ICCV2021] Safety-aware Motion Prediction with Unseen Vehicles for Autonomous Driving

Local Similarity Pattern and Cost Self-Reassembling for Deep Stereo Matching Networks

REBEL: Relation Extraction By End-to-end Language generation

Unified unsupervised and semi-supervised domain adaptation network for cross-scenario face anti-spoofing, Pattern Recognition

[NeurIPS 2021] Introspective Distillation for Robust Question Answering

Additional environments compatible with OpenAI gym

Course content and resources for the AIAIART course.

Anti-UAV base on PaddleDetection

torchlm is aims to build a high level pipeline for face landmarks detection, it supports training, evaluating, exporting, inference(Python/C++) and 100+ data augmentations

A PyTorch library and evaluation platform for end-to-end compression research

source code and pre-trained/fine-tuned checkpoint for NAACL 2021 paper LightningDOT

A repository for benchmarking neural vocoders by their quality and speed.

DNA sequence classification by Deep Neural Network