MoViNet-pytorch

Pytorch unofficial implementation of MoViNets: Mobile Video Networks for Efficient Video Recognition.
Authors: Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, Boqing Gong (Google Research)
[Authors' Implementation]

Stream Buffer

Clean stream buffer

It is required to clean the buffer after all the clips of the same video have been processed.

model.clean_activation_buffers()

Usage

Click on "Open in Colab" to open an example of training on HMDB-51

installation

pip install git+https://github.com/Atze00/MoViNet-pytorch.git

How to build a model

Use causal = True to use the model with stream buffer, causal = False will use standard convolutions

from movinets import MoViNet
from movinets.config import _C

MoViNetA0 = MoViNet(_C.MODEL.MoViNetA0, causal = True, pretrained = True )
MoViNetA1 = MoViNet(_C.MODEL.MoViNetA1, causal = True, pretrained = True )
...

Load weights

Use pretrained = True to use the model with pretrained weights

    """
    If pretrained is True:
        num_classes is set to 600,
        conv_type is set to "3d" if causal is False, "2plus1d" if causal is True
        tf_like is set to True
    """
model = MoViNet(_C.MODEL.MoViNetA0, causal = True, pretrained = True )
model = MoViNet(_C.MODEL.MoViNetA0, causal = False, pretrained = True )

Training loop examples

Training loop with stream buffer

def train_iter(model, optimz, data_load, n_clips = 5, n_clip_frames=8):
    """
    In causal mode with stream buffer a single video is fed to the network
    using subclips of lenght n_clip_frames. 
    n_clips*n_clip_frames should be equal to the total number of frames presents
    in the video.
    
    n_clips : number of clips that are used
    n_clip_frames : number of frame contained in each clip
    """
    
    #clean the buffer of activations
    model.clean_activation_buffers()
    optimz.zero_grad()
    for i, data, target in enumerate(data_load):
        #backward pass for each clip
        for j in range(n_clips):
          out = F.log_softmax(model(data[:,:,(n_clip_frames)*(j):(n_clip_frames)*(j+1)]), dim=1)
          loss = F.nll_loss(out, target)/n_clips
          loss.backward()
        optimz.step()
        optimz.zero_grad()
        
        #clean the buffer of activations
        model.clean_activation_buffers()

Training loop with standard convolutions

def train_iter(model, optimz, data_load):

    optimz.zero_grad()
    for i, (data,_ , target) in enumerate(data_load):
        out = F.log_softmax(model(data), dim=1)
        loss = F.nll_loss(out, target)
        loss.backward()
        optimz.step()
        optimz.zero_grad()

Pretrained models

Weights

The weights are loaded from the tensorflow models released by the authors, trained on kinetics.

Base Models

Base models implement standard 3D convolutions without stream buffers.

Model Name	Top-1 Accuracy*	Top-5 Accuracy*	Input Shape
MoViNet-A0-Base	72.28	90.92	50 x 172 x 172
MoViNet-A1-Base	76.69	93.40	50 x 172 x 172
MoViNet-A2-Base	78.62	94.17	50 x 224 x 224
MoViNet-A3-Base	81.79	95.67	120 x 256 x 256
MoViNet-A4-Base	83.48	96.16	80 x 290 x 290
MoViNet-A5-Base	84.27	96.39	120 x 320 x 320

Model Name	Top-1 Accuracy*	Top-5 Accuracy*	Input Shape**
MoViNet-A0-Stream	72.05	90.63	50 x 172 x 172
MoViNet-A1-Stream	76.45	93.25	50 x 172 x 172
MoViNet-A2-Stream	78.40	94.05	50 x 224 x 224

**In streaming mode, the number of frames correspond to the total accumulated duration of the 10-second clip.

*Accuracy reported on the official repository for the dataset kinetics 600, It has not been tested by me. It should be the same since the tf models and the reimplemented pytorch models output the same results [Test].

I currently haven't tested the speed of the streaming models, feel free to test and contribute.

Status

Currently are available the pretrained models for the following architectures:

I currently have no plans to include streaming version of A3,A4,A5. Those models are too slow for most mobile applications.

Testing

I recommend to create a new environment for testing and run the following command to install all the required packages:
pip install -r tests/test_requirements.txt

Citations

@article{kondratyuk2021movinets,
  title={MoViNets: Mobile Video Networks for Efficient Video Recognition},
  author={Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Matthew Brown, and Boqing Gong},
  journal={arXiv preprint arXiv:2103.11511},
  year={2021}
}

MoViNets PyTorch implementation: Mobile Video Networks for Efficient Video Recognition;

Related tags

Overview

MoViNet-pytorch

Stream Buffer

Clean stream buffer

Usage

installation

How to build a model

Load weights

Training loop examples

Pretrained models

Weights

Base Models

Status

Testing

Citations

Owner

Python script to download the celebA-HQ dataset from google drive

ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation.

Code release for Universal Domain Adaptation(CVPR 2019)

Reference models and tools for Cloud TPUs.

The official start-up code for paper "FFA-IR: Towards an Explainable and Reliable Medical Report Generation Benchmark."

pyspark🍒🥭 is delicious，just eat it!😋😋

Efficient training of deep recommenders on cloud.

Deep Learning for 3D Point Clouds: A Survey (IEEE TPAMI, 2020)

Reimplementation of Dynamic Multi-scale filters for Semantic Segmentation.

Text-to-Image generation

NudeNet: Neural Nets for Nudity Classification, Detection and selective censoring

Self-Supervised Document-to-Document Similarity Ranking via Contextualized Language Models and Hierarchical Inference

A distributed, plug-n-play algorithm for multi-robot applications with a priori non-computable objective functions

Towards Boosting the Accuracy of Non-Latin Scene Text Recognition

A library that allows for inference on probabilistic models

MATLAB codes of the book "Digital Image Processing Fourth Edition" converted to Python

[3DV 2021] Channel-Wise Attention-Based Network for Self-Supervised Monocular Depth Estimation

Using LSTM write Tang poetry

This repository contains the data and code for the paper "Diverse Text Generation via Variational Encoder-Decoder Models with Gaussian Process Priors" ([email protected])

A clear, concise, simple yet powerful and efficient API for deep learning.