This repo contains the pytorch implementation for Dynamic Concept Learner (accepted by ICLR 2021).

Last update: Jan 06, 2023

Related tags

Deep Learning DCL-Release

Overview

DCL-PyTorch

Pytorch implementation for the Dynamic Concept Learner (DCL). More details can be found at the project page.

Framework

Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning
Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee K. Wong, Joshua B. Tenenbaum, and Chuang Gan

Prerequisites

Python 3
PyTorch 1.0 or higher, with NVIDIA CUDA Support
Other required python packages specified by requirements.txt. See the Installation.

Installation

Install Jacinle: Clone the package, and add the bin path to your global PATH environment variable:

git clone https://github.com/vacancy/Jacinle --recursive
export PATH=<path_to_jacinle>/bin:$PATH

Clone this repository:

git clone https://github.com/zfchenUnique/DCL-Release.git --recursive

Create a conda environment for NS-CL, and install the requirements. This includes the required python packages from both Jacinle NS-CL. Most of the required packages have been included in the built-in anaconda package:

Dataset preparation

Download videos, video annotation, questions and answers, and object proposals accordingly from the official website
Transform videos into ".png" frames with ffmpeg.

Organize the data as shown below.

clevrer
├── annotation_00000-01000
│   ├── annotation_00000.json
│   ├── annotation_00001.json
│   └── ...
├── ...
├── image_00000-01000
│   │   ├── 1.png
│   │   ├── 2.png
│   │   └── ...
│   └── ...
├── ...
├── questions
│   ├── train.json
│   ├── validation.json
│   └── test.json
├── proposals
│   ├── proposal_00000.json
│   ├── proposal_00001.json
│   └── ...

Fast Evaluation

Download the extracted object trajectories from google drive.
Git clone the dynamic model, download image proposals and the pretrained propNet models and make dynamic prediction by

    git clone https://github.com/zfchenUnique/clevrer_dynamic_propnet.git
    cd clevrer_dynamic_propnet
    sh ./scripts/eval_fast_release_v2.sh 0

Download the pretrained DCL model and parsed programs.

   sh scripts/script_test_prp_clevrer_qa.sh 0

Get the accuracy on evalAI.

Step-by-step Training

Step 1: download the proposals from the region proposal network and extract object trajectories for train and val set by

   sh scripts/script_gen_tubes.sh

Step 2: train a concept learner with descriptive and explanatory questions for static concepts (i.e. color, shape and material)

   sh scripts/script_train_dcl_stage1.sh 0

Step 3: extract static attributes & refine object trajectories extract static attributes

   sh scripts/script_extract_attribute.sh

refine object trajectories

   sh scripts/script_gen_tubes_refine.sh

Step 4: extract predictive and counterfactual scenes by

    cd clevrer_dynamic_propnet
    sh ./scripts/train_tube_box_only.sh # train
    sh ./scripts/train_tube.sh # train
    sh ./scripts/eval_fast_release_v2.sh 0 # val

Step 5: train DCL with all questions and the refined trajectories

   sh scripts/script_train_dcl_stage2.sh 0

Generalization to CLEVRER-Grounding

Step 1: download expression annotation and parsed programs from google drive
Step 2: evaluate the performance on CLEVRER-Grounding

    sh ./scripts/script_grounding.sh  0
    jac-crun 0 scripts/script_evaluate_grounding.py

Generalization to CLEVRER-Retrieval

Step 1: download expression annotation and parsed programs from google drive
Step 2: evaluate the performance on CLEVRER-Retrieval

    sh ./scripts/script_retrieval.sh  0
    jac-crun 0 scripts/script_evaluate_retrieval.py

Extension to Tower Blocks

Step 1: download question annotation and videos from google drive
Step 2: train on Tower block QA

    sh ./scripts/script_train_blocks.sh 0

Step 3: download the pretrain model from google drive and evaluate on Tower block QA

    sh ./scripts/script_eval_blocks.sh 0

Others

Citation

If you find this repo useful in your research, please consider citing:

@inproceedings{zfchen2021iclr,
    title={Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning},
    author={Chen, Zhenfang and Mao, Jiayuan and Wu, Jiajun and Wong, Kwan-Yee K and Tenenbaum, Joshua B. and Gan, Chuang},
    booktitle={International Conference on Learning Representations},
    year={2021}
    }

This repo contains the pytorch implementation for Dynamic Concept Learner (accepted by ICLR 2021).

Related tags

Overview

DCL-PyTorch

Framework

Prerequisites

Installation

Dataset preparation

Fast Evaluation

Step-by-step Training

Generalization to CLEVRER-Grounding

Generalization to CLEVRER-Retrieval

Extension to Tower Blocks

Others

Citation

Owner

Zhenfang Chen

A Strong Baseline for Image Semantic Segmentation

A pytorch-based real-time segmentation model for autonomous driving

Out-of-Domain Human Mesh Reconstruction via Dynamic Bilevel Online Adaptation

Source code for paper "ATP: AMRize Than Parse! Enhancing AMR Parsing with PseudoAMRs" @NAACL-2022

Freecodecamp Scientific Computing with Python Certification; Solution for Challenge 2: Time Calculator

UCSD Oasis platform

Code for ICLR2018 paper: Improving GAN Training via Binarized Representation Entropy (BRE) Regularization - Y. Cao · W Ding · Y.C. Lui · R. Huang

Source code of the paper "Deep Learning of Latent Variable Models for Industrial Process Monitoring".

PyTorch IPFS Dataset

A Data Annotation Tool for Semantic Segmentation, Object Detection and Lane Line Detection.(In Development Stage)

Single-Shot Motion Completion with Transformer

Text-to-Music Retrieval using Pre-defined/Data-driven Emotion Embeddings

Few-shot Neural Architecture Search

Official implementation of Deep Convolutional Dictionary Learning for Image Denoising.

Generative Handwriting using LSTM Mixture Density Network with TensorFlow

Detector for Log4Shell exploitation attempts

CNN Based Meta-Learning for Noisy Image Classification and Template Matching

An implementation of the proximal policy optimization algorithm

Code to generate datasets used in "How Useful is Self-Supervised Pretraining for Visual Tasks?"

Simple PyTorch hierarchical models.