PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision.

Last update: Sep 14, 2022

Related tags

Overview

PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision

@misc{CV2018,
  author =       {Donny You ([email protected])},
  howpublished = {\url{https://github.com/donnyyou/PyTorchCV}},
  year =         {2018}
}

This repository provides source code for some deep learning based cv problems. We'll do our best to keep this repository up to date. If you do find a problem about this repository, please raise it as an issue. We will fix it immediately.

Implemented Papers

Image Classification
- VGG: Very Deep Convolutional Networks for Large-Scale Image Recognition
- ResNet: Deep Residual Learning for Image Recognition
- DenseNet: Densely Connected Convolutional Networks
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- ShuffleNet V2: Practical Guidelines for Ecient CNN Architecture Design
Semantic Segmentation
- DeepLabV3: Rethinking Atrous Convolution for Semantic Image Segmentation
- PSPNet: Pyramid Scene Parsing Network
- DenseASPP: DenseASPP for Semantic Segmentation in Street Scenes
Object Detection
- SSD: Single Shot MultiBox Detector
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- YOLOv3: An Incremental Improvement
- FPN: Feature Pyramid Networks for Object Detection
Pose Estimation
- CPM: Convolutional Pose Machines
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
Instance Segmentation
- Mask R-CNN

Performances with PyTorchCV

Image Classification

ResNet: Deep Residual Learning for Image Recognition

Semantic Segmentation

PSPNet: Pyramid Scene Parsing Network

Model	Backbone	Training data	Testing data	mIOU	Pixel Acc	Setting
PSPNet Origin	3x3-ResNet101	ADE20K train	ADE20K val	41.96	80.64	-
PSPNet Ours	7x7-ResNet101	ADE20K train	ADE20K val	44.18	80.91	PSPNet

Object Detection

SSD: Single Shot MultiBox Detector

Model	Backbone	Training data	Testing data	mAP	FPS	Setting
SSD-300 Origin	VGG16	VOC07+12 trainval	VOC07 test	0.772	-	-
SSD-300 Ours	VGG16	VOC07+12 trainval	VOC07 test	0.786	-	SSD300
SSD-512 Origin	VGG16	VOC07+12 trainval	VOC07 test	0.798	-	-
SSD-512 Ours	VGG16	VOC07+12 trainval	VOC07 test	0.808	-	SSD512

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Model	Backbone	Training data	Testing data	mAP	FPS	Setting
Faster R-CNN Origin	VGG16	VOC07 trainval	VOC07 test	0.699	-	-
Faster R-CNN Ours	VGG16	VOC07 trainval	VOC07 test	0.706	-	Faster R-CNN

YOLOv3: An Incremental Improvement

Pose Estimation

OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

Instance Segmentation

Mask R-CNN

Commands with PyTorchCV

Take PSPNet as an example. ("tag" could be any string, include an empty one.)

Training

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh train tag

Resume Training

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh train tag

Validate

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh val tag

Testing:

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh test tag

Examples with PyTorchCV

Example output of VGG19-OpenPose

PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision.

Related tags

Overview

PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision

Implemented Papers

Performances with PyTorchCV

Image Classification

Semantic Segmentation

Object Detection

Pose Estimation

Instance Segmentation

Commands with PyTorchCV

Examples with PyTorchCV

Owner

Donny You

NVIDIA container runtime

Beyond imagenet attack (accepted by ICLR 2022) towards crafting adversarial examples for black-box domains.

DCGAN LSGAN WGAN-GP DRAGAN PyTorch

BASH - Biomechanical Animated Skinned Human

Share a benchmark that can easily apply reinforcement learning in Job-shop-scheduling

Official Pytorch Implementation of 3DV2021 paper: SAFA: Structure Aware Face Animation.

The codes reproduce the figures and statistics in the paper, "Controlling for multiple covariates," by Mark Tygert.

Official Implementation for "StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery" (ICCV 2021 Oral)

Pmapper is a super-resolution and deconvolution toolkit for python 3.6+

A set of Deep Reinforcement Learning Agents implemented in Tensorflow.

Convolutional Neural Networks

RobustART: Benchmarking Robustness on Architecture Design and Training Techniques

code for our paper "Source Data-absent Unsupervised Domain Adaptation through Hypothesis Transfer and Labeling Transfer"

Deep-Learning-Image-Captioning - Implementing convolutional and recurrent neural networks in Keras to generate sentence descriptions of images

Anomaly detection analysis and labeling tool, specifically for multiple time series (one time series per category)

A module that used for encrypt code which includes RSA and AES

Rename Images with Auto Generated Neural Image Captions

Official Repository for the ICCV 2021 paper "PixelSynth: Generating a 3D-Consistent Experience from a Single Image"

Trading and Backtesting environment for training reinforcement learning agent or simple rule base algo.

Predicting Auction Sale Price using the kaggle bulldozer auction sales data: Modeling with Ensembles vs Neural Network