Clockwork Convnets for Video Semantic Segmentation

This is the reference implementation of arxiv:1608.03609:

Clockwork Convnets for Video Semantic Segmentation
Evan Shelhamer*, Kate Rakelly*, Judy Hoffman*, Trevor Darrell
arXiv:1605.06211

This project reproduces results from the arxiv and demonstrates how to execute staged fully convolutional networks (FCNs) on video in Caffe by controlling the net through the Python interface. In this way this these experiments are a proof-of-concept implementation of clockwork, and further development is needed to achieve peak efficiency (such as pre-fetching video data layers, threshold GPU layers, and a native Caffe library edition of the staged forward pass for pipelining).

For simple reference, refer to these (display only) editions of the experiments:

Cityscapes Clockwork
YouTube Frame Differencing
YouTube Clockwork
YouTube Pipelining
Synthetic PASCAL VOC Video
Dataset Walkthroughs for YouTube, NYUDv2, and Cityscapes

Contents

notebooks: interactive code and documentation that carries out the experiments (in jupyter/ipython format).
nets: the net specification of the various FCNs in this work, and the pre-trained weights (see installation instructions).
caffe: the Caffe framework, included as a git submodule pointing to a compatible version
datasets: input-output for PASCAL VOC, NYUDv2, YouTube-Objects, and Cityscapes
lib: helpers for executing networks, scoring metrics, and plotting

License

This project is licensed for open non-commercial distribution under the UC Regents license; see LICENSE. Its dependencies, such as Caffe, are subject to their own respective licenses.

Requirements & Installation

Caffe, Python, and Jupyter are necessary for all of the experiments. Any installation or general Caffe inquiries should be directed to the caffe-users mailing list.

Install Caffe. See the installation guide and try Caffe through Docker (recommended). Make sure to configure pycaffe, the Caffe Python interface, too.
Install Python, and then install our required packages listed in requirements.txt. For instance, for x in $(cat requirements.txt); do pip install $x; done should do.
Install Jupyter, the interface for viewing, executing, and altering the notebooks.
Configure your PYTHONPATH as indicated by the included .envrc so that this project dir and pycaffe are included.
Download the model weights for this project and place them in nets.

Now you can explore the notebooks by firing up Jupyter.

Clockwork Convnets for Video Semantic Segmentation

Related tags

Overview

Clockwork Convnets for Video Semantic Segmentation

License

Requirements & Installation

Owner

Evan Shelhamer

Logsig-RNN: a novel network for robust and efficient skeleton-based action recognition

[CVPR 2022 Oral] Crafting Better Contrastive Views for Siamese Representation Learning

GE2340 project source code without credentials.

Towards Open-World Feature Extrapolation: An Inductive Graph Learning Approach

A package related to building quasi-fibration symmetries

NOMAD - A blackbox optimization software

This is the repository for paper NEEDLE: Towards Non-invertible Backdoor Attack to Deep Learning Models.

QA-GNN: Question Answering using Language Models and Knowledge Graphs

Building a real-time environment using webcam frame division in OpenCV and classify cropped images using a fine-tuned vision transformers on hybryd datasets samples for facial emotion recognition.

Use your Philips Hue lights as Racing Flags. Works with Assetto Corsa, Assetto Corsa Competizione and iRacing.

This is a repository with the code for the ACL 2019 paper

A library for efficient similarity search and clustering of dense vectors.

Block Sparse movement pruning

PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.

An AFL implementation with UnTracer (our coverage-guided tracer)

Distance Encoding for GNN Design

Reference models and tools for Cloud TPUs.

Autonomous Perception: 3D Object Detection with Complex-YOLO

A python library for highly configurable transformers - easing model architecture search and experimentation.

The official TensorFlow implementation of the paper Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition