Includes PyTorch -> Keras model porting code for ConvNeXt family of models with fine-tuning and inference notebooks.

Last update: Dec 06, 2022

Overview

ConvNeXt-TF

This repository provides TensorFlow / Keras implementations of different ConvNeXt [1] variants. It also provides the TensorFlow / Keras models that have been populated with the original ConvNeXt pre-trained weights available from [2]. These models are not blackbox SavedModels i.e., they can be fully expanded into tf.keras.Model objects and one can call all the utility functions on them (example: .summary()).

As of today, all the TensorFlow / Keras variants of the models listed here are available in this repository except for the isotropic ones. This list includes the ImageNet-1k as well as ImageNet-21k models.

Refer to the "Using the models" section to get started. Additionally, here's a related blog post that jots down my experience.

Conversion

TensorFlow / Keras implementations are available in models/convnext_tf.py. Conversion utilities are in convert.py.

Models

The converted models are available on TF-Hub.

There should be a total of 15 different models each having two variants: classifier and feature extractor. You can load any model and get started like so:

import tensorflow as tf

model_gcs_path = "gs://tfhub-modules/sayakpaul/convnext_tiny_1k_224/1/uncompressed"
model = tf.keras.models.load_model(model_gcs_path)
print(model.summary(expand_nested=True))

The model names are interpreted as follows:

convnext_large_21k_1k_384: This means that the model was first pre-trained on the ImageNet-21k dataset and was then fine-tuned on the ImageNet-1k dataset. Resolution used during pre-training and fine-tuning: 384x384. large denotes the topology of the underlying model.
convnext_large_1k_224: Means that the model was pre-trained on the ImageNet-1k dataset with a resolution of 224x224.

Results

Results are on ImageNet-1k validation set (top-1 accuracy).

name	original [email protected]	keras [email protected]
convnext_tiny_1k_224	82.1	81.312
convnext_small_1k_224	83.1	82.392
convnext_base_1k_224	83.8	83.28
convnext_base_1k_384	85.1	84.876
convnext_large_1k_224	84.3	83.844
convnext_large_1k_384	85.5	85.376

convnext_base_21k_1k_224	85.8	85.364
convnext_base_21k_1k_384	86.8	86.79
convnext_large_21k_1k_224	86.6	86.36
convnext_large_21k_1k_384	87.5	87.504
convnext_xlarge_21k_1k_224	87.0	86.732
convnext_xlarge_21k_1k_384	87.8	87.68

Differences in the results are primarily because of the differences in the library implementations especially how image resizing is implemented in PyTorch and TensorFlow. Results can be verified with the code in i1k_eval. Logs are available at this URL.

Using the models

Pre-trained models:

Off-the-shelf classification: Colab Notebook
Fine-tuning: Colab Notebook

Randomly initialized models:

from models.convnext_tf import get_convnext_model

convnext_tiny = get_convnext_model()
print(convnext_tiny.summary(expand_nested=True))

To view different model configurations, refer here.

Upcoming (contributions welcome)

Align layer initializers (useful if someone wanted to train the models from scratch)
Allow the models to accept arbitrary shapes (useful for downstream tasks)
Convert the isotropic models as well
Fine-tuning notebook (thanks to awsaf49)
Off-the-shelf-classification notebook
Publish models on TF-Hub

References

[1] ConvNeXt paper: https://arxiv.org/abs/2201.03545

[2] Official ConvNeXt code: https://github.com/facebookresearch/ConvNeXt

Includes PyTorch -> Keras model porting code for ConvNeXt family of models with fine-tuning and inference notebooks.

Related tags

Overview

ConvNeXt-TF

Conversion

Models

Results

Using the models

Upcoming (contributions welcome)

References

Acknowledgements

Owner

Sayak Paul

Hysterese plugin with two temperature offset areas

Research shows Google collects 20x more data from Android than Apple collects from iOS. Block this non-consensual telemetry using pihole blocklists.

OneShot Learning-based hotword detection.

Picasso: A CUDA-based Library for Deep Learning over 3D Meshes

(CVPR2021) Kaleido-BERT: Vision-Language Pre-training on Fashion Domain

Cache Requests in Deta Bases and Echo them with Deta Micros

Pytorch implementation of FlowNet by Dosovitskiy et al.

Code for the paper "Controllable Video Captioning with an Exemplar Sentence"

BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Code for ICE-BeeM paper - NeurIPS 2020

The Malware Open-source Threat Intelligence Family dataset contains 3,095 disarmed PE malware samples from 454 families

BasicNeuralNetwork - This project looks over the basic structure of a neural network and how machine learning training algorithms work

[CVPR 2022] Deep Equilibrium Optical Flow Estimation

Vehicle speed detection with python

Explainer for black box models that predict molecule properties

Official pytorch implementation for Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion (CVPR 2022)

[ICCV2021] IICNet: A Generic Framework for Reversible Image Conversion

A brand new hub for Scene Graph Generation methods based on MMdetection (2021). The pipeline of from detection, scene graph generation to downstream tasks (e.g., image cpationing) is supported. Pytorch version implementation of HetH (ECCV 2020) and TopicSG (ICCV 2021) is included.

This is the repository for Learning to Generate Piano Music With Sustain Pedals

Official code for 'Pixel-wise Energy-biased Abstention Learning for Anomaly Segmentationon Complex Urban Driving Scenes'