PyTorch implementation for 3D human pose estimation

Last update: Dec 22, 2022

Related tags

Overview

Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach

This repository is the PyTorch implementation for the network presented in:

Xingyi Zhou, Qixing Huang, Xiao Sun, Xiangyang Xue, Yichen Wei, Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach ICCV 2017 (arXiv:1704.02447)

Note: This repository has been updated and is different from the method discribed in the paper. To fully reproduce the results in the paper, please checkout the original torch implementation or our pytorch re-implementation branch (slightly worse than torch). We also provide a clean 2D hourglass network branch.

The updates include:

Change network backbone to ResNet50 with deconvolution layers (Xiao et al. ECCV2018). Training is now about 3x faster than the original hourglass net backbone (but no significant performance improvement).
Change the depth regression sub-network to a one-layer depth map (described in our StarMap project).
Change the Human3.6M dataset to official release in ECCV18 challenge.
Update from python 2.7 and pytorch 0.1.12 to python 3.6 and pytorch 0.4.1.

Contact: [email protected]

Installation

The code was tested with Anaconda Python 3.6 and PyTorch v0.4.1. After install Anaconda and Pytorch:

Clone the repo:

POSE_ROOT=/path/to/clone/pytorch-pose-hg-3d
git clone https://github.com/xingyizhou/pytorch-pose-hg-3d POSE_ROOT

Install dependencies (opencv, and progressbar):

conda install --channel https://conda.anaconda.org/menpo opencv
conda install --channel https://conda.anaconda.org/auto progress

Disable cudnn for batch_norm (see issue):

# PYTORCH=/path/to/pytorch
# for pytorch v0.4.0
sed -i "1194s/torch\.backends\.cudnn\.enabled/False/g" ${PYTORCH}/torch/nn/functional.py
# for pytorch v0.4.1
sed -i "1254s/torch\.backends\.cudnn\.enabled/False/g" ${PYTORCH}/torch/nn/functional.py

Optionally, install tensorboard for visializing training.
```
pip install tensorflow
```

Demo

Download our pre-trained model and move it to models.
Run python demo.py --demo /path/to/image/or/image/folder [--gpus -1] [--load_model /path/to/model].

--gpus -1 is for CPU mode. We provide example images in images/. For testing your own image, it is important that the person should be at the center of the image and most of the body parts should be within the image.

Benchmark Testing

To test our model on Human3.6 dataset run

python main.py --exp_id test --task human3d --dataset fusion_3d --load_model ../models/fusion_3d_var.pth --test --full_test

The expected results should be 64.55mm.

Training

Prepare the training data:

Download images from MPII dataset and their annotation in json format (train.json and val.json) (from Xiao et al. ECCV2018).
Download Human3.6M ECCV challenge dataset.
Download meta data (2D bounding box) of the Human3.6 dataset (from Sun et al. ECCV 2018).
Place the data (or create symlinks) to make the data folder like:

${POSE_ROOT}
|-- data
`-- |-- mpii
    `-- |-- annot
        |   |-- train.json
        |   |-- valid.json
        `-- images
            |-- 000001163.jpg
            |-- 000003072.jpg
`-- |-- h36m
    `-- |-- ECCV18_Challenge
        |   |-- Train
        |   |-- Val
        `-- msra_cache
            `-- |-- HM36_eccv_challenge_Train_cache
                |   |-- HM36_eccv_challenge_Train_w288xh384_keypoint_jnt_bbox_db.pkl
                `-- HM36_eccv_challenge_Val_cache
                    |-- HM36_eccv_challenge_Val_w288xh384_keypoint_jnt_bbox_db.pkl

Stage1: Train 2D pose only. model, log

python main.py --exp_id mpii

Stage2: Train on 2D and 3D data without geometry loss (drop LR at 45 epochs). model, log

python main.py --exp_id fusion_3d --task human3d --dataset fusion_3d --ratio_3d 1 --weight_3d 0.1 --load_model ../exp/mpii/model_last.pth --num_epoch 60 --lr_step 45

Stage3: Train with geometry loss. model, log

python main.py --exp_id fusion_3d_var --task human3d --dataset fusion_3d --ratio_3d 1 --weight_3d 0.1 --weight_var 0.01 --load_model ../models/fusion_3d.pth  --num_epoch 10 --lr 1e-4

Citation

@InProceedings{Zhou_2017_ICCV,
author = {Zhou, Xingyi and Huang, Qixing and Sun, Xiao and Xue, Xiangyang and Wei, Yichen},
title = {Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach},
booktitle = {The IEEE International Conference on Computer Vision (ICCV)},
month = {Oct},
year = {2017}
}

PyTorch implementation for 3D human pose estimation

Related tags

Overview

Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach

Installation

Demo

Benchmark Testing

Training

Citation

Owner

Xingyi Zhou

Python Classes: Medical Insurance Project using Object Oriented Programming Concepts

An Open Source Machine Learning Framework for Everyone

Fully Adaptive Bayesian Algorithm for Data Analysis (FABADA) is a new approach of noise reduction methods. In this repository is shown the package developed for this new method based on \citepaper.

Self-Supervised Pre-Training for Transformer-Based Person Re-Identification

Proposal, Tracking and Segmentation (PTS): A Cascaded Network for Video Object Segmentation

DeepStruc is a Conditional Variational Autoencoder which can predict the mono-metallic nanoparticle from a Pair Distribution Function.

LightNet++: Boosted Light-weighted Networks for Real-time Semantic Segmentation

Frigate - NVR With Realtime Object Detection for IP Cameras

SOLO and SOLOv2 for instance segmentation, ECCV 2020 & NeurIPS 2020.

PyTorch code for our paper "Image Super-Resolution with Non-Local Sparse Attention" (CVPR2021).

Code for One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning (AAAI 2022)

Permeability Prediction Via Multi Scale 3D CNN

A simple, fast, and efficient object detector without FPN

CLASP - Contrastive Language-Aminoacid Sequence Pretraining

Code for: Gradient-based Hierarchical Clustering using Continuous Representations of Trees in Hyperbolic Space. Nicholas Monath, Manzil Zaheer, Daniel Silva, Andrew McCallum, Amr Ahmed. KDD 2019.

Spherical CNNs

GeoTransformer - Geometric Transformer for Fast and Robust Point Cloud Registration

Repo for parser tensorflow(.pb) and tflite(.tflite)

Python PID Tuner - Based on a FOPDT model obtained using a Open Loop Process Reaction Curve

Official repo for QHack—the quantum machine learning hackathon