Neural Dynamic Policies for End-to-End Sensorimotor Learning

Last update: Dec 11, 2022

Related tags

Deep Learning neural-dynamic-policies

Overview

Neural Dynamic Policies for End-to-End Sensorimotor Learning

In NeurIPS 2020 (Spotlight) [Project Website] [Project Video]

Shikhar Bahl, Mustafa Mukadam, Abhinav Gupta, Deepak Pathak
Carnegie Mellon University & Facebook AI Research

This is a PyTorch based implementation for our NeurIPS 2020 paper on Neural Dynamic Policies for end-to-end sensorimotor learning. In this work, we begin to close this gap and embed dynamics structure into deep neural network-based policies by reparameterizing action spaces with differential equations. We propose Neural Dynamic Policies (NDPs) that make predictions in trajectory distribution space as opposed to prior policy learning methods where action represents the raw control space. The embedded structure allow us to perform end-to-end policy learning under both reinforcement and imitation learning setups. If you find this work useful in your research, please cite:

  @inproceedings{bahl2020neural,
    Author = { Bahl, Shikhar and Mukadam, Mustafa and
    Gupta, Abhinav and Pathak, Deepak},
    Title = {Neural Dynamic Policies for End-to-End Sensorimotor Learning},
    Booktitle = {NeurIPS},
    Year = {2020}
  }

1) Installation and Usage

This code is based on PyTorch. This code needs MuJoCo 1.5 to run. To install and setup the code, run the following commands:

#create directory for data and add dependencies
cd neural-dynamic-polices; mkdir data/
git clone https://github.com/rll/rllab.git
git clone https://github.com/openai/baselines.git

#create virtual env
conda create --name ndp python=3.5
source activate ndp

#install requirements
pip install -r requirements.txt
#OR try
conda env create -f ndp.yaml

Training imitation learning

cd neural-dynamic-polices
# name of the experiment
python main_il.py --name NAME

Training RL: run the script run_rl.sh. ENV_NAME is the environment (could be throw, pick, push, soccer, faucet). ALGO-TYPE is the algorithm (dmp for NDPs, ppo for PPO [Schulman et al., 2017] and ppo-multi for the multistep actor-critic architecture we present in our paper).

sh run_rl.sh ENV_NAME ALGO-TYPE EXP_ID SEED

In order to visualize trained models/policies, use the same exact arguments as used for training but call vis_policy.sh

  sh vis_policy.sh ENV_NAME ALGO-TYPE EXP_ID SEED

2) Other helpful pointers

3) Acknowledgements

Neural Dynamic Policies for End-to-End Sensorimotor Learning

Related tags

Overview

Neural Dynamic Policies for End-to-End Sensorimotor Learning

In NeurIPS 2020 (Spotlight) [Project Website] [Project Video]

1) Installation and Usage

2) Other helpful pointers

3) Acknowledgements

Owner

Shikhar Bahl

PyTorch-centric library for evaluating and enhancing the robustness of AI technologies

Western-3DSlicer-Modules - Point-Set Registrations for Ultrasound Probe Calibrations

magiCARP: Contrastive Authoring+Reviewing Pretraining

Code, Data and Demo for Paper: Controllable Generation from Pre-trained Language Models via Inverse Prompting

Virtual Dance Reality Stage is a feature that offers you to share a stage with another user virtually.

A higher performance pytorch implementation of DeepLab V3 Plus(DeepLab v3+)

Source code for Fathony, Sahu, Willmott, & Kolter, "Multiplicative Filter Networks", ICLR 2021.

A pure PyTorch batched computation implementation of "CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition"

alfred-py: A deep learning utility library for human

Pytorch implementation of paper "Learning Co-segmentation by Segment Swapping for Retrieval and Discovery"

Implementation of the Remixer Block from the Remixer paper, in Pytorch

NAVER BoostCamp Final Project

Code for the prototype tool in our paper "CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data Poisoning".

This is an easy python software which allows to sort images with faces by gender and after by age.

code for generating data set ES-ImageNet with corresponding training code

face property detection pytorch

Data manipulation and transformation for audio signal processing, powered by PyTorch

FcaNet: Frequency Channel Attention Networks

Homepage of paper: Paint Transformer: Feed Forward Neural Painting with Stroke Prediction, ICCV 2021.

Rethinking Nearest Neighbors for Visual Classification

Neural Dynamic Policies for End-to-End Sensorimotor Learning

Related tags

Overview

Neural Dynamic Policies for End-to-End Sensorimotor Learning

In NeurIPS 2020 (Spotlight) [Project Website] [Project Video]

1) Installation and Usage

2) Other helpful pointers

3) Acknowledgements

Owner

Shikhar Bahl

PyTorch-centric library for evaluating and enhancing the robustness of AI technologies

Western-3DSlicer-Modules - Point-Set Registrations for Ultrasound Probe Calibrations

magiCARP: Contrastive Authoring+Reviewing Pretraining

Code, Data and Demo for Paper: Controllable Generation from Pre-trained Language Models via Inverse Prompting

Virtual Dance Reality Stage is a feature that offers you to share a stage with another user virtually.

A higher performance pytorch implementation of DeepLab V3 Plus(DeepLab v3+)

Source code for Fathony, Sahu, Willmott, & Kolter, "Multiplicative Filter Networks", ICLR 2021.

A pure PyTorch batched computation implementation of "CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition"

alfred-py: A deep learning utility library for **human**

Pytorch implementation of paper "Learning Co-segmentation by Segment Swapping for Retrieval and Discovery"

Implementation of the Remixer Block from the Remixer paper, in Pytorch

NAVER BoostCamp Final Project

Code for the prototype tool in our paper "CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data Poisoning".

This is an easy python software which allows to sort images with faces by gender and after by age.

code for generating data set ES-ImageNet with corresponding training code

face property detection pytorch

Data manipulation and transformation for audio signal processing, powered by PyTorch

FcaNet: Frequency Channel Attention Networks

Homepage of paper: Paint Transformer: Feed Forward Neural Painting with Stroke Prediction, ICCV 2021.

Rethinking Nearest Neighbors for Visual Classification

alfred-py: A deep learning utility library for human