Neural Dynamic Policies for End-to-End Sensorimotor Learning

Last update: Dec 11, 2022

Related tags

Deep Learning neural-dynamic-policies

Overview

Neural Dynamic Policies for End-to-End Sensorimotor Learning

In NeurIPS 2020 (Spotlight) [Project Website] [Project Video]

Shikhar Bahl, Mustafa Mukadam, Abhinav Gupta, Deepak Pathak
Carnegie Mellon University & Facebook AI Research

This is a PyTorch based implementation for our NeurIPS 2020 paper on Neural Dynamic Policies for end-to-end sensorimotor learning. In this work, we begin to close this gap and embed dynamics structure into deep neural network-based policies by reparameterizing action spaces with differential equations. We propose Neural Dynamic Policies (NDPs) that make predictions in trajectory distribution space as opposed to prior policy learning methods where action represents the raw control space. The embedded structure allow us to perform end-to-end policy learning under both reinforcement and imitation learning setups. If you find this work useful in your research, please cite:

  @inproceedings{bahl2020neural,
    Author = { Bahl, Shikhar and Mukadam, Mustafa and
    Gupta, Abhinav and Pathak, Deepak},
    Title = {Neural Dynamic Policies for End-to-End Sensorimotor Learning},
    Booktitle = {NeurIPS},
    Year = {2020}
  }

1) Installation and Usage

This code is based on PyTorch. This code needs MuJoCo 1.5 to run. To install and setup the code, run the following commands:

#create directory for data and add dependencies
cd neural-dynamic-polices; mkdir data/
git clone https://github.com/rll/rllab.git
git clone https://github.com/openai/baselines.git

#create virtual env
conda create --name ndp python=3.5
source activate ndp

#install requirements
pip install -r requirements.txt
#OR try
conda env create -f ndp.yaml

Training imitation learning

cd neural-dynamic-polices
# name of the experiment
python main_il.py --name NAME

Training RL: run the script run_rl.sh. ENV_NAME is the environment (could be throw, pick, push, soccer, faucet). ALGO-TYPE is the algorithm (dmp for NDPs, ppo for PPO [Schulman et al., 2017] and ppo-multi for the multistep actor-critic architecture we present in our paper).

sh run_rl.sh ENV_NAME ALGO-TYPE EXP_ID SEED

In order to visualize trained models/policies, use the same exact arguments as used for training but call vis_policy.sh

  sh vis_policy.sh ENV_NAME ALGO-TYPE EXP_ID SEED

2) Other helpful pointers

3) Acknowledgements

Neural Dynamic Policies for End-to-End Sensorimotor Learning

Related tags

Overview

Neural Dynamic Policies for End-to-End Sensorimotor Learning

In NeurIPS 2020 (Spotlight) [Project Website] [Project Video]

1) Installation and Usage

2) Other helpful pointers

3) Acknowledgements

Owner

Shikhar Bahl

Algorithmic Trading using RNN

Mixed Transformer UNet for Medical Image Segmentation

A memory-efficient implementation of DenseNets

PyTorch implementation of Lip to Speech Synthesis with Visual Context Attentional GAN (NeurIPS2021)

RAANet: Range-Aware Attention Network for LiDAR-based 3D Object Detection with Auxiliary Density Level Estimation

Single/multi view image(s) to voxel reconstruction using a recurrent neural network

Modified prey-predator system - Modified prey–predator model describes the rate of change for each species by adding coupling terms.

PyTorch Implementation of PIXOR: Real-time 3D Object Detection from Point Clouds

Continual Learning of Electronic Health Records (EHR).

Model-based 3D Hand Reconstruction via Self-Supervised Learning, CVPR2021

Semantic Segmentation Architectures Implemented in PyTorch

Escaping the Gradient Vanishing: Periodic Alternatives of Softmax in Attention Mechanism

A very impractical 3D rendering engine that runs in the python terminal.

Image data augmentation scheduler for albumentations transforms

Learning to Self-Train for Semi-Supervised Few-Shot

🥇 LG-AI-Challenge 2022 1위 솔루션 입니다.

Implementation of paper: "Image Super-Resolution Using Dense Skip Connections" in PyTorch

Event-forecasting - Event Forecasting Algorithms With Python

This game was designed to encourage young people not to gamble on lotteries, as the probablity of correctly guessing the number is infinitesimal!

Repository for the paper "Exploring the Sensory Spaces of English Perceptual Verbs in Natural Language Data"