This project is a re-implementation of MASTER: Multi-Aspect Non-local Network for Scene Text Recognition by MMOCR

Last update: Nov 17, 2022

Related tags

Overview

MASTER-mmocr

About The Project
- Dependency
Getting Started
- Prerequisites
- Installation
Usage
Result
Coming Soon
License
Citations
Acknowledgements

About The Project

This project is a re-implementation of MASTER: Multi-Aspect Non-local Network for Scene Text Recognition by MMOCR，which is an open-source toolbox based on PyTorch. The overall architecture will be shown below.

Dependency

Getting Started

Prerequisites

Use Synthetic image datasets: SynthText (Synth800k), MJSynth (Synth90k) for training.
Real image datasets: IIIT5K, SVT, IC03, IC13, IC15, SVTP, CUTE80 for testing.

Dataset download link.
Change dataset path in MASTER config.

Installation

Install mmdetection. click here for details.

# We embed mmdetection-2.11.0 source code into this project.
# You can cd and install it (recommend).
cd ./mmdetection-2.11.0
pip install -v -e .

Install mmocr. click here for details.

# install mmocr
cd ./MASTER_mmocr
pip install -v -e .

Install mmcv-full-1.3.4. click here for details.

pip install mmcv-full=={mmcv_version} -f https://download.openmmlab.com/mmcv/dist/{cu_version}/{torch_version}/index.html

# install mmcv-full-1.3.4 with torch version 1.8.0 cuda_version 10.2
pip install mmcv-full==1.3.4 -f https://download.openmmlab.com/mmcv/dist/cu102/torch1.8.0/index.html

Usage

The usage of this project, is consistent with MMOCR-0.2.0. You can click here for mmocr usage details.

For training, run command

CUDA_VISIBLE_DEVICES={device_id} PORT={port_number} ./tools/dist_train.sh {config_path} {work_dir} {gpu_number}

# example
CUDA_VISIBLE_DEVICES=0 PORT=29500 ./tools/dist_train.sh ./configs/textrecog/master/master_ResnetExtra_academic_dataset_dynamic_mmfp16.py /expr/mmocr_text_line_recognition/ 1

PS :

As mentioned in Prerequisites part, we use synthetic image datasets for training and real image datasets for evalutating. The 7 real image datasets mentioned above will be evaluated at each evaluation interval.

Result

Dataset	Paper reported accuracy	Our accuracy
IIIT5K	95.0	95.07
SVT	90.6	90.42
IC03	96.4	95.58
IC13	95.3	96.03
IC15	79.4	80.95
SVTP	84.5	84.34
CUTE80	87.5	90.62

Coming Soon

1st Solution for ICDAR 2021 Competition on Scientific Table Image Recognition to Latex.

License

This project is licensed under the MIT License. See LICENSE for more details.

Citations

If you find MASTER useful please cite paper:

@article{Lu2021MASTER,
  title={{MASTER}: Multi-Aspect Non-local Network for Scene Text Recognition},
  author={Ning Lu and Wenwen Yu and Xianbiao Qi and Yihao Chen and Ping Gong and Rong Xiao and Xiang Bai},
  journal={Pattern Recognition},
  year={2021}
}

This project is a re-implementation of MASTER: Multi-Aspect Non-local Network for Scene Text Recognition by MMOCR

Related tags

Overview

MASTER-mmocr

Contents

About The Project

Dependency

Getting Started

Prerequisites

Installation

Usage

Result

Coming Soon

License

Citations

Acknowledgements

Owner

Jianquan Ye

DeepCO3: Deep Instance Co-segmentation by Co-peak Search and Co-saliency

Learn about Spice.ai with in-depth samples

A Number Recognition algorithm

YuNetのPythonでのONNX、TensorFlow-Lite推論サンプル

Transformers based fully on MLPs

Exploring Image Deblurring via Blur Kernel Space (CVPR'21)

MPI-IS Mesh Processing Library

TAUFE: Task-Agnostic Undesirable Feature DeactivationUsing Out-of-Distribution Data

NeurIPS 2021 Datasets and Benchmarks Track

MultiLexNorm 2021 competition system from ÚFAL

Implementation of Kronecker Attention in Pytorch

[CVPR 2021] Unsupervised 3D Shape Completion through GAN Inversion

Hierarchical Metadata-Aware Document Categorization under Weak Supervision (WSDM'21)

MoViNets PyTorch implementation: Mobile Video Networks for Efficient Video Recognition;

Scalable Optical Flow-based Image Montaging and Alignment

We evaluate our method on different datasets (including ShapeNet, CUB-200-2011, and Pascal3D+) and achieve state-of-the-art results, outperforming all the other supervised and unsupervised methods and 3D representations, all in terms of performance, accuracy, and training time.

System Design course at HSE (2021)

DiscoNet: Learning Distilled Collaboration Graph for Multi-Agent Perception [NeurIPS 2021]

ONNX Runtime Web demo is an interactive demo portal showing real use cases running ONNX Runtime Web in VueJS.