ML-Decoder: Scalable and Versatile Classification Head

Last update: Jan 04, 2023

Related tags

Deep Learning ML_Decoder

Overview

ML-Decoder: Scalable and Versatile Classification Head

Paper

Official PyTorch Implementation

Tal Ridnik, Gilad Sharir, Avi Ben-Cohen, Emanuel Ben-Baruch, Asaf Noy
DAMO Academy, Alibaba Group

Abstract

In this paper, we introduce ML-Decoder, a new attention-based classification head. ML-Decoder predicts the existence of class labels via queries, and enables better utilization of spatial data compared to global average pooling. By redesigning the decoder architecture, and using a novel group-decoding scheme, ML-Decoder is highly efficient, and can scale well to thousands of classes. Compared to using a larger backbone, ML-Decoder consistently provides a better speed-accuracy trade-off. ML-Decoder is also versatile - it can be used as a drop-in replacement for various classification heads, and generalize to unseen classes when operated with word queries. Novel query augmentations further improve its generalization ability. Using ML-Decoder, we achieve state-of-the-art results on several classification tasks: on MS-COCO multi-label, we reach 91.4% mAP; on NUS-WIDE zero-shot, we reach 31.1% ZSL mAP; and on ImageNet single-label, we reach with vanilla ResNet50 backbone a new top score of 80.7%, without extra data or distillation.

ML-Decoder Implementation

ML-Decoder implementation is available here. It can be easily integrated into any backbone using this example code:

ml_decoder_head = MLDecoder(num_classes) # initilization

spatial_embeddings = self.backbone(input_image) # backbone generates spatial embeddings      
 
logits = ml_decoder_head(spatial_embeddings) # transfrom spatial embeddings to logits

Training Code

We will share a full reproduction code for the article results.

Multi-label Training Code

A reproduction code for MS-COCO multi-label:

python train.py  \
--data=/home/datasets/coco2014/ \
--model_name=tresnet_l \
--image_size=448

Single-label Training Code

Our single-label training code uses the excellent timm repo. Reproduction code is currently from a fork, we will work toward a full merge to the main repo.

git clone https://github.com/mrT23/pytorch-image-models.git

This is the code for A2 configuration training, with ML-Decoder (--use-ml-decoder-head=1):

python -u -m torch.distributed.launch --nproc_per_node=8 \
--nnodes=1 \
--node_rank=0 \
./train.py \
/data/imagenet/ \
--amp \
-b=256 \
--epochs=300 \
--drop-path=0.05 \
--opt=lamb \
--weight-decay=0.02 \
--sched='cosine' \
--lr=4e-3 \
--warmup-epochs=5 \
--model=resnet50 \
--aa=rand-m7-mstd0.5-inc1 \
--reprob=0.0 \
--remode='pixel' \
--mixup=0.1 \
--cutmix=1.0 \
--aug-repeats 3 \
--bce-target-thresh 0.2 \
--smoothing=0 \
--bce-loss \
--train-interpolation=bicubic \
--use-ml-decoder-head=1

ZSL Training Code

Reproduction code for ZSL is WIP.

Citation

@misc{ridnik2021mldecoder,
      title={ML-Decoder: Scalable and Versatile Classification Head}, 
      author={Tal Ridnik and Gilad Sharir and Avi Ben-Cohen and Emanuel Ben-Baruch and Asaf Noy},
      year={2021},
      eprint={2111.12933},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

ML-Decoder: Scalable and Versatile Classification Head

Related tags

Overview

ML-Decoder: Scalable and Versatile Classification Head

ML-Decoder Implementation

Training Code

Multi-label Training Code

Single-label Training Code

ZSL Training Code

Citation

Owner

Custom implementation of Corrleation Module

CLIP: Connecting Text and Image (Learning Transferable Visual Models From Natural Language Supervision)

DeepGNN is a framework for training machine learning models on large scale graph data.

BT-Unet: A-Self-supervised-learning-framework-for-biomedical-image-segmentation-using-Barlow-Twins

A project that uses optical flow and machine learning to detect aimhacking in video clips.

First-Order Probabilistic Programming Language

Implementation of FitVid video prediction model in JAX/Flax.

Lux AI environment interface for RLlib multi-agents

Continual learning with sketched Jacobian approximations

Real-time Neural Representation Fusion for Robust Volumetric Mapping

A Domain-Agnostic Benchmark for Self-Supervised Learning

上海交通大学全自动抢课脚本，支持准点开抢与抢课后持续捡漏两种模式。2021/06/08更新。

TensorFlow 101: Introduction to Deep Learning for Python Within TensorFlow

Dirty Pixels: Towards End-to-End Image Processing and Perception

Simple improvement of VQVAE that allow to generate x2 sized images compared to baseline

Code for Robust Contrastive Learning against Noisy Views

TensorFlow implementation of Elastic Weight Consolidation

This project generates news headlines using a Long Short-Term Memory (LSTM) neural network.

This is my research project for the Irving Center for Cancer Dynamics/Azizi Lab, Columbia University.

Optimising chemical reactions using machine learning