textspotter - An End-to-End TextSpotter with Explicit Alignment and Attention

Last update: Nov 10, 2022

Overview

An End-to-End TextSpotter with Explicit Alignment and Attention

This is initially described in our CVPR 2018 paper.

Getting Started

Installation

Clone the code

git clone https://github.com/tonghe90/textspotter
cd textspotter

Install caffe. You can follow this this tutorial. If you have build problem about std::allocater, please refer to this #3

# make sure you set WITH_PYTHON_LAYER := 1
# change Makefile.config according to your library path
cp Makefile.config.example Makefile.config
make clean
make -j8
make pycaffe

Training

we provide part of the training code. But you can not run this directly. 
We have give the comment in the [train.pt](https://github.com/tonghe90/textspotter/models/train.pt).
You have to write your own layer, IOUloss layer. We cannot publish this for some IP reason. 
To be noticed: 
[L6902](https://github.com/tonghe90/textspotter/models/train.pt#L6902) 
[L6947](https://github.com/tonghe90/textspotter/models/train.pt#L6907)

Testing

install editdistance and pyclipper: pip install editdistance and pip install pyclipper
After Caffe is set up, you need to download a trained model (about 40M) from Google Drive. This model is trained with VGG800k and finetuned on ICDAR2015.
Run python test.py --img=./imgs/img_105.jpg
hyperparameters:

cfg.py --mean_val ==> mean value during the testing.
       --max_len ==> maximum length of the text string (here we take 25, meaning a word can contain 25 characters at most.)
       --recog_th ==> the threshold during the recognition process. The score for a word is the average mean of every character.
       --word_score ==> the threshold for those words that contain number or symbols for they are not contained in the dictionary.

test.py --weight ==> weights file of caffemodel
        --prototxt-iou ==> the prototxt file for detection.
        --prototxt-lstm ==> the prototxt file for recognition.
        --img ==> the folder or img file for testing. The format can be added in ./pylayer/tool is_image function.
        --scales-ms ==> multiscales input for input during the testing process.
        --thresholds-ms ==> corresponding thresholds of text region for multiscale inputs.
        --nms ==> nms threshold for testing
        --save-dir ==> the dir for save results in format of ICDAR2015 submition.

One thing should be noted: the recognition results are achieved by comparing direct output with words in dictionary, which has about 90K lexicons. 
These lexicons don't contain any number and symbol. You can delete dictionary reference part and directly output recognition results.

Citation

If you use this code for your research, please cite our papers.

@inproceedings{tong2018,
  title={An End-to-End TextSpotter with Explicit Alignment and Attention},
  author={T. He and Z. Tian and W. Huang and C. Shen and Y. Qiao and C. Sun},
  booktitle={Computer Vision and Pattern Recognition (CVPR), 2018 IEEE Conference on},
  year={2018}
}

License

This code is for NON-COMMERCIAL purposes only. For commerical purposes, please contact Chunhua Shen [email protected]. This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, version 3. Please refer to http://www.gnu.org/licenses/ for more details.

textspotter - An End-to-End TextSpotter with Explicit Alignment and Attention

Related tags

Overview

An End-to-End TextSpotter with Explicit Alignment and Attention

Getting Started

Installation

Training

Testing

Citation

License

Owner

Tong He

Pixel art search engine for opengameart

Official implementation of "An Image is Worth 16x16 Words, What is a Video Worth?" (2021 paper)

When Age-Invariant Face Recognition Meets Face Age Synthesis: A Multi-Task Learning Framework (CVPR 2021 oral)

Repository for Scene Text Detection with Supervised Pyramid Context Network with tensorflow.

code for our ICCV 2021 paper "DeepCAD: A Deep Generative Network for Computer-Aided Design Models"

Single Shot Text Detector with Regional Attention

Code for the head detector (HeadHunter) proposed in our CVPR 2021 paper Tracking Pedestrian Heads in Dense Crowd.

Handwritten Text Recognition (HTR) system implemented with TensorFlow (TF) and trained on the IAM off-line HTR dataset. This Neural Network (NN) model recognizes the text contained in the images of segmented words.

CRAFT-Pyotorch：Character Region Awareness for Text Detection Reimplementation for Pytorch

~1000 book pages + OpenCV + python = page regions identified as paragraphs, lines, images, captions, etc.

A real-time dolly zoom camera effect

Face Detection with DLIB

OCR powered screen-capture tool to capture information instead of images

A Tensorflow model for text recognition (CNN + seq2seq with visual attention) available as a Python package and compatible with Google Cloud ML Engine.

A python programusing Tkinter graphics library to randomize questions and answers contained in text files

SCOUTER: Slot Attention-based Classifier for Explainable Image Recognition

Detect and fix skew in images containing text

Let's explore how we can extract text from forms

Semantic-based Patch Detection for Binary Programs

Brief idea about our project is mentioned in project presentation file.