Optical Character Recognition + Instance Segmentation for russian and english languages

Last update: Dec 19, 2022

Overview

Распознавание рукописного текста в школьных тетрадях

Соревнование, проводимое в рамках олимпиады НТО, разработанное Сбером. Платформа ODS.

Результаты Public

Задача

Вам нужно разработать алгоритм, который способен распознать рукописный текст в школьных тетрадях. В качестве входных данных вам будут предоставлены фотографии целых листов. Предсказание модели — список распознанных строк с координатами полигонов и получившимся текстом.

Как должно работать решение?

Последовательность двух моделей: сегментации и распознавания. Сначала сегментационная модель предсказывает полигоны маски каждого слова на фото. Затем эти слова вырезаются из изображения по контуру маски (получаются кропы на каждое слово) и подаются в модель распознавания. В итоге получается список распознанных слов с их координатами.

Модели

Instance Segmentation

модель X101-FPN из зоопарка моделей detectron2 + аугментации + высокое разрешение

Optical Character Recognition (OCR)

архитектура CRNN с бекбоном Resnet-34, предобученным на топ 1 модели соревнования Digital Peter

Beam Search

модель KenLM, обученная на данных сорвенования Feedback, Решу ОГЭ/ЕГЭ, а также CTCDecoder

Ресурсы & Submit

Christofari с NVIDIA Tesla V100 и образом jupyter-cuda10.1-tf2.3.0-pt1.6.0-gpu:0.0.82

Мы не гарантируем поддержку сабмита всё время, поэтому предоставляем 2 ссылки: Google Drive и Yandex

Цитирование

@misc{nto-ai-text-recognition,
  author =       {Arseniy Shahmatov and Gerasomiv Maxim},
  title =        {notebook-recognition},
  howpublished = {\url{https://github.com/Lednik7/nto-ai-text-recognition}},
  year =         {2022}
}

You might also like...

Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow

Mask R-CNN for Object Detection and Segmentation This is an implementation of Mask R-CNN on Python 3, Keras, and TensorFlow. The model generates bound

22.5k Jan 4, 2023

This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows" on Object Detection and Instance Segmentation.

Swin Transformer for Object Detection This repo contains the supported code and configuration files to reproduce object detection results of Swin Tran

1.4k Dec 30, 2022

Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch.

Faster R-CNN and Mask R-CNN in PyTorch 1.0 maskrcnn-benchmark has been deprecated. Please see detectron2, which includes implementations for all model

9k Jan 4, 2023

Object detection and instance segmentation toolkit based on PaddlePaddle.

9.3k Jan 2, 2023

DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision

The Official PyTorch Implementation of DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision

3 Oct 15, 2021

The PyTorch implementation of DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision.

DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision The PyTorch implementation of DiscoBox: Weakly Supe

1 Oct 23, 2021

Optical Character Recognition + Instance Segmentation for russian and english languages

Related tags

Overview

Распознавание рукописного текста в школьных тетрадях

Соревнование, проводимое в рамках олимпиады НТО, разработанное Сбером. Платформа ODS.

Результаты Public

Задача

Как должно работать решение?

Модели

Ресурсы & Submit

Цитирование

You might also like...

Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow

This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows" on Object Detection and Instance Segmentation.

Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch.

Object detection and instance segmentation toolkit based on PaddlePaddle.

DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision

The PyTorch implementation of DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision.

Numbering permanent and deciduous teeth via deep instance segmentation in panoramic X-rays

Res2Net for Instance segmentation and Object detection using MaskRCNN

Keras implementation of PersonLab for Multi-Person Pose Estimation and Instance Segmentation.

Releases(v1.0.0)

v1.0.0(Mar 6, 2022)

Owner

Gerasimov Maxim

TensorFlow 2 AI/ML library wrapper for openFrameworks

Efficient Conformer: Progressive Downsampling and Grouped Attention for Automatic Speech Recognition

Code for CVPR 2018 paper --- Texture Mapping for 3D Reconstruction with RGB-D Sensor

Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with Transformers

FLVIS: Feedback Loop Based Visual Initial SLAM

This repository implements Douzero's interface to IGCA.

Submodular Subset Selection for Active Domain Adaptation (ICCV 2021)

Autolfads-tf2 - A TensorFlow 2.0 implementation of Latent Factor Analysis via Dynamical Systems (LFADS) and AutoLFADS

Learning and Building Convolutional Neural Networks using PyTorch

GDSC-ML Team Interview Task

MemStream: Memory-Based Anomaly Detection in Multi-Aspect Streams with Concept Drift

Instance-level Image Retrieval using Reranking Transformers

Official implementation of the ICCV 2021 paper "Joint Inductive and Transductive Learning for Video Object Segmentation"

A simple baseline for 3d human pose estimation in tensorflow. Presented at ICCV 17.

Pairwise learning neural link prediction for ogb link prediction

CCPD: a diverse and well-annotated dataset for license plate detection and recognition

Categorizing comments on YouTube into different categories.

Tool for working with Y-chromosome data from YFull and FTDNA

Code repository for the work "Multi-Domain Incremental Learning for Semantic Segmentation", accepted at WACV 2022

A MNIST-like fashion product database. Benchmark