A TensorFlow implementation of FCN-8s

Last update: Aug 08, 2022

Overview

FCN-8s implementation in TensorFlow

Overview
Examples and demo video
Dependencies
How to use it
Download pre-trained VGG-16

Overview

This is a TensorFlow implementation of the FCN-8s model architecture for semantic image segmentation introduced by Shelhamer et al. in the paper Fully Convolutional Networks for Semantic Segmentation.

This repository only contains the 'all-at-once' version of the FCN-8s model, which converges significantly faster than the version trained in stages. A convolutionalized VGG-16 model trained on ImageNet classification is provided and serves as the encoder of the FCN-8s. Sufficient documentation and a tutorial on how to train, evaluate and use the model for prediction are also provided. Some useful TensorBoard summaries can be recorded out of the box.

Examples and demo video

Below are some prediction examples of the model trained on the Cityscapes dataset for 13,000 steps at batch size 16, at which point the model achieves a mean IoU of 38.2% on the validation dataset. This is far from convergence of course, the purpose of these examples is just to demonstrate that the code works and the model learns. You can watch the model in action on the Cityscapes demo videos here.

Dependencies

Python 3.x
TensorFlow 1.x
Numpy
Scipy
OpenCV (for data augmentation)
tqdm

How to use it

fcn8s_tutorial.ipynb explains how to train and evaluate the model and how to make and visualize predictions.

Download pre-trained VGG-16

You can download the pre-trained, convolutionalized VGG-16 model here

A TensorFlow implementation of FCN-8s

Related tags

Overview

FCN-8s implementation in TensorFlow

Contents

Overview

Examples and demo video

Dependencies

How to use it

Download pre-trained VGG-16

Owner

Pierluigi Ferrari

DziriBERT: a Pre-trained Language Model for the Algerian Dialect

Code for the paper "TadGAN: Time Series Anomaly Detection Using Generative Adversarial Networks"

SubOmiEmbed: Self-supervised Representation Learning of Multi-omics Data for Cancer Type Classification

Pytorch implementation of the paper DocEnTr: An End-to-End Document Image Enhancement Transformer.

Brain Tumor Detection with Tensorflow Neural Networks.

Dynamic hair modeling from monocular videos using deep neural networks

CM-NAS: Cross-Modality Neural Architecture Search for Visible-Infrared Person Re-Identification (ICCV2021)

RodoSol-ALPR Dataset

Embeddinghub is a database built for machine learning embeddings.

Lightweight, Python library for fast and reproducible experimentation :microscope:

ParmeSan: Sanitizer-guided Greybox Fuzzing

Explaining Hyperparameter Optimization via PDPs

code and models for "Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation"

A repo with study material, exercises, examples, etc for Devnet SPAUTO

The implementation of PEMP in paper "Prior-Enhanced Few-Shot Segmentation with Meta-Prototypes"

IEEE Winter Conference on Applications of Computer Vision 2022 Accepted

StyleSwin: Transformer-based GAN for High-resolution Image Generation

[CVPR 2021] Generative Hierarchical Features from Synthesizing Images

RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation

The PyTorch improved version of TPAMI 2017 paper: Face Alignment in Full Pose Range: A 3D Total Solution.