Python library to extract tabular data from images and scanned PDFs

Last update: Dec 31, 2022

Overview

ExtractTable - API to extract tabular data from images and scanned PDFs

The motivation is to make it easy for developers to extract tabular data from images or scanned PDF files without worrying about the table area, column coordinates, rotation et al.

Prerequisite

API Key: All requests to ExtractTable are authorized by an API Key. FREE credits here. The same API Key can also be used for conversions on the browser at Web Pro.

Installation

pip install -U ExtractTable

Basic Usage

Ok, enough selling. Let the ease in coding do the talk, and the output encourages you to buy credits; put that timer on and count the LOC.

from ExtractTable import ExtractTable
et_sess = ExtractTable(api_key=YOUR_API_KEY)        # Replace your VALID API Key here
print(et_sess.check_usage())        # Checks the API Key validity as well as shows associated plan usage 
table_data = et_sess.process_file(filepath=Location_of_Image_with_Tables, output_format="df")

# To process PDF, make use of pages ("1", "1,3-4", "all") params in the read_pdf function
table_data = et_sess.process_file(filepath=Location_of_PDF_with_Tables, output_format="df", pages="all")

Detailed Library Usage

The tutorial available at takes you through

1. Installation
2. Import and check version
3. Create Session & Validate API Key
    3.1 Create Session with your API Key
    3.2 Validate the Key and check the plan usage
    3.3 Check Usage Details
4. Trigger the extraction process
    4.1 Accepted Input Types
    4.2 Process an IMAGE Input
    4.3 Process a PDF Input
    4.4 Output options
    4.5 Explore session objects
5. Explore the Output
    5.1 Output Structure
    5.2 Output Details
6. Make Corrections
    6.1 Split Merged Rows
    6.2 Split Merged Columns
    6.3 Fix Decimal Format
    6.4 Fix Date Format
7. Helpful Code Snippets
    7.1 Get text data
    7.2 Table output to Excel

Woahh, as simple as that ?!

Certainly. Do you know the current ExtractTable users use it for

Bank Statement
Medical Records
Invoice Details
Tax forms
Tender Notices

Its up to you now to explore the ways.

Explore

check the complete server response of the latest job with et_sess.ServerResponse.json()

{
    "JobStatus": <string>,                              # Status of the triggered Process  @ JOB-LEVEL
    "Pages": <integer>,                                 # Number of pages processed in this request @ PAGE-LEVEL
    "Tables": [<list of key-value objects of table>     # List of all tables found @ TABLE-LEVEL
        {
            "Page": <integer>,                              ## Page number in which this table is found
            "CharacterConfidence": <float>,                 ## Accuracy of Characters recognized from the input-page
            "LayoutConfidence": <float>,                    ## Accuracy of table layout's design decision
            "TableJson": <dict>,                            ## Table Cell Text in key-value format with index orientation - {row#: {col#: <str>}}
            "TableCoordinates": <dict>,                     ## Top-left & Bottom-right Cell Coordinates - {row#: {col#: <list(x1,y1,x2,y2)>}}
            "TableConfidence": <dict>                       ## Cell level accuracy of detected characters - {row#: {col#: <float>}}
        },
    {...}                                               ## ... more "Tables" objects
    ],
    "Lines": [<list of key-value objects>               # Pagewise Line details @ PAGE-LEVEL
        {
            "Page": <integer>,                          # Page number in which the lines are found
            "CharacterConfidence": <float>,             # Average Accuracy of all Characters recognized from the input-page
            "LinesArray": [
                <list of key-value objects of line>     # Ordered list of lines in this page @ LINE-LEVEL
                {
                    "Line": <str>,                          ## Detected text of the complete line
                    "WordsArray": [
                        <list of key-value objects>         ## Word level datails in this line @ WORD-LEVEL
                        {
                            "Conf": <float>,                    ### Accuracy of recognized characters of the word
                            "Word": <str>,                      ### Detected text of the word
                            "Loc": [x1, y1, x2, y2]             ### Top-left & Bottom-right coordinates, w.r.t the input-page width-height dimensions
                        },
                    {...}                                   ### More "WordsArray" objects
                    ]
                },
            {...}                                       ## More "LinesArray" objects
            ]
        },
    {...}                                               # More Pagewise "Lines" details
    ]
}

Bug Reports

Bug reports/fixes are most welcome and greatly appreciated with API credits. For support reach us at [email protected]

License

This project is licensed under the Apache License 2.0, see the LICENSE file for details.

Social Media

Comments

bug: holding when the program running after some samples

Describe the bug A clear and concise description of what the bug is. keep holding my apI key prefix is o6No6aqYRhrQ2MWxtDDyTeHiiUg****

To Reproduce Steps to reproduce the behavior: or the code you tried

Expected behavior A clear and concise description of what you expected to happen.

Additional context Add any other context about the problem here.
bug

opened by franztao 5
bug: function "et_sess.save_output(output_folder, output_format="csv")" output file, the file name lack some alpha of the origin full name

Describe the bug A clear and concise description of what the bug is. my picture name is all suffix png. such as "[email protected]_14-1-4.png"

To Reproduce Steps to reproduce the behavior: or the code you tried

Expected behavior A clear and concise description of what you expected to happen.

Additional context Add any other context about the problem here.
bug

opened by franztao 3
found some bugs and list the bugs out

Describe the bug A clear and concise description of what the bug is. 1.不能识别出垮列的文本，识别成表格时，不符合逻辑的分开成两边

2.不能识别加减号,can not recognize Plus minus sign. 31.2 + 4.98 3.不能够识别上下标，can not recognize subscript and supscript. 4.ocr识别丢失字符 loss some recognized tokens 5.长的表格，有部分没有识别出来 long size table,can not recognize the bottem part 6.cell中有化学式的，识别不出来,when there is chemical formulate in cell, can not recognize the table

To Reproduce Steps to reproduce the behavior: or the code you tried

Expected behavior A clear and concise description of what you expected to happen. I can solve these problems with us.

Additional context Add any other context about the problem here.
bug

opened by franztao 2
question: what meaning is LayoutConfidence?

"CharacterConfidence": , # Average Accuracy of all Characters recognized from the input-page "LayoutConfidence": , ## Accuracy of table layout's design decision please give out the detaild decription or calculate function code about CharacterConfidence,LayoutConfidence
good first issue

opened by franztao 2
Invalid cross-device link
Describe the bug On some OS, we can not save output file to temporary directory (let's say /tmp) and move it to a new place. It throws the following error :

os.replace(each_tbl_path, os.path.join(output_folder, input_fname+os.path.basename(each_tbl_path))) OSError: [Errno 18] Invalid cross-device link: '/tmp/tmp7hqcm0fh/_table_1.csv' -> '/var/www/python/app/tmp/details_table_1.csv'

After checking the source code, it appears ExtractTable use os.replace to move the file. This method does not support moving file from a partition to an other : https://stackoverflow.com/questions/42392600/oserror-errno-18-invalid-cross-device-link

To Reproduce I use Python 3.6 in a venv. You will need two different system parts, and invoke save_output from ExtractTable-py library, to save file from a filesystem to an other. I have not tried, but I think you can simply reproduce this bug by invoking os.replace without calling ExtractTable-py.

Expected behavior Move the file from a filesystem to an other. I think using shutil.move would be a preferable way to achieve file moving than os.replace.
bug
opened by Elegye 2
MakeCorrections API - How do you chain corrections

Hi there, I'm trying to use multiple correction commands but it isn't working as the object becomes a list after the first correction. Is there something I'm missing here? Thanks!
good first issue

opened by kylebutts 1
character ocr can support latex format?

Is your feature request related to a problem? Please describe. A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Describe alternatives you've considered A clear and concise description of any alternative solutions or features you've considered.

Describe the solution you'd like [optional, but helpful] A clear and concise description of what you want to happen.

Additional context Add any other context or screenshots about the feature request here.

opened by franztao 1
please, do you have tools of transform ExtracTable output file type to CoCo file type(other open source Detection file type)?

Is your feature request related to a problem? Please describe. A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Describe alternatives you've considered A clear and concise description of any alternative solutions or features you've considered.

Describe the solution you'd like [optional, but helpful] A clear and concise description of what you want to happen.

Additional context Add any other context or screenshots about the feature request here.

opened by franztao 1
Custom output path when the output_format is csv

Is your feature request related to a problem? Please describe. When the output_format is set to csv the csv file is written to some random path in /tmp location.

Describe the solution you'd like [optional, but helpful] Define a parameter in the process_file like output_file which takes the absolute path where the file needs to be written along with the file name

opened by padmano 1
Is it possible to get the data in excel by maintaining table structure?

Is your feature request related to a problem? Please describe. A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Describe alternatives you've considered A clear and concise description of any alternative solutions or features you've considered.

Describe the solution you'd like [optional, but helpful] A clear and concise description of what you want to happen.

Additional context Add any other context or screenshots about the feature request here.

opened by jcthink 1

Character and Layout Confidence

Hi, need some definition material for Character and Layout Confidence like how it is calculated mathematically using below code. Thanks.

for idx, each_table in enumerate(et_sess.ServerResponse.json()['Tables']):
    print("CharacterConfidence = ", each_table['CharacterConfidence'])
    print("LayoutConfidence = ", each_table['LayoutConfidence'])

good first issue

opened by muhdzubair 1

Consider user hints on the table structure information

Is your feature request related to a problem? Please describe. "while you do whatever you want, why not consider the our hints" is the developers feedback on many instances

Describe alternatives you've considered Developers are tackling with their custom post processing.

Describe the solution you'd like [optional, but helpful] Pros: May be it is a worth taking a look as most of the post processing involves in similar approaches that resolves majority issues. Cons: computing cost
feature/idea

opened by akshowhini 0
Capture Vertically center aligned columns

Refer: https://stackoverflow.com/questions/58238981/extracting-table-from-a-pdf-table-without-vertical-lines

Do not miss: Joelgeraci's comment to the question
feature/idea

opened by akshowhini 0

Releases(v2.4.0)

v2.4.0(Jul 18, 2022)

Use corrections.save_output() to save the output to a folder
Source code(tar.gz)
Source code(zip)
ExtractTable-2.4.0-py3-none-any.whl(18.90 KB)
ExtractTable-2.4.0.tar.gz(16.16 KB)
v2.3.1(May 6, 2022)
Fix processing splitted PDFs

Support downloading BigFile

Source code(tar.gz)
Source code(zip)
ExtractTable-2.3.1-py3-none-any.whl(18.38 KB)
ExtractTable-2.3.1.tar.gz(16.01 KB)
v2.2.0(Apr 20, 2021)
View your transactions processed in the last 24 hours

Give user the ability to make character error corrections

Source code(tar.gz)
Source code(zip)
ExtractTable-2.2.0-py3-none-any.whl(20.52 KB)
v2.1.2(Nov 6, 2020)

To provide user control on whether to output the row & column numbers in the output file
Source code(tar.gz)
Source code(zip)
ExtractTable-2.1.2-py3-none-any.whl(20.13 KB)
ExtractTable-2.1.2.tar.gz(12.04 KB)
v2.1.0(Aug 27, 2020)
Data Cleaning on the server output made easy with MakeCorrections class.

split_merged_rows

split_merged_columns

fix_decimal_format

fix_date_format functionalities

added server_response attribute to the session for easy reference

Update Google Colab Tutorial

Save tables to multiple sheets of a single excel file

save_output functionality in session to save Tables & Text output to local

Updated Tutorial in example-code.ipynb
Source code(tar.gz)
Source code(zip)
ExtractTable-2.1.0-py3-none-any.whl(20.12 KB)
ExtractTable-2.1.0.tar.gz(12.03 KB)
v2.0.2(Jul 4, 2020)

To handle Invalid Object Exception when processing big files
Source code(tar.gz)
Source code(zip)
ExtractTable-2.0.2-py3-none-any.whl(17.28 KB)
ExtractTable-2.0.2.tar.gz(9.27 KB)
v2.0.1(Jul 3, 2020)
#28

Maintain column and row indices order

Display JobId & Wait message for async transactions

Source code(tar.gz)
Source code(zip)
ExtractTable-2.0.1-py3-none-any.whl(17.27 KB)
ExtractTable-2.0.1.tar.gz(9.27 KB)
v2.0.0(Apr 30, 2020)

To support the below added features at API level • Tables + Text Data: non-tabular text along with the tabular data (when tables undetected, by default, response gets text data) • Text Accuracy Details: page level character accuracy details • Cell / Word Coordinates: x,y coordinates of all words and table's cell data • Cell / Word Level Accuracy: word level accuracy details • Non-English characters: for non-english alphabets like Mandrin, Japanese etc
Source code(tar.gz)
Source code(zip)
ExtractTable-2.0.0-py3-none-any.whl(17.04 KB)
ExtractTable-2.0.0.tar.gz(9.09 KB)
v1.2.1.2(Dec 1, 2019)

Fixed Columns were not in order in the output
Source code(tar.gz)
Source code(zip)
ExtractTable-1.2.1.2-py3-none-any.whl(14.09 KB)
ExtractTable-1.2.1.2.tar.gz(8.01 KB)
v1.1.0(Oct 20, 2019)

Support URL as input which downloads the file to the temporary directory for the instance, auto deleted on success
Source code(tar.gz)
Source code(zip)
v1.0.1(Oct 7, 2019)

The first stable library to make it easier for python developers to use ExtractTable's API to extract tabular data (table) from images and scanned PDFs without worrying about table area, column regions, image rotation et al.
Source code(tar.gz)
Source code(zip)
ExtractTable-1.0.1-py3-none-any.whl(12.42 KB)
ExtractTable-1.0.1.tar.gz(6.54 KB)

Owner

Org. Account

You, I and they have the same problem to solve !?!?

GitHub Repository https://extracttable.com

This tool will help you convert your text to handwriting xD

So your teacher asked you to upload written assignments? Hate writing assigments? This tool will help you convert your text to handwriting xD

4.2k Jan 07, 2023

Single Shot Text Detector with Regional Attention

Single Shot Text Detector with Regional Attention Introduction SSTD is initially described in our ICCV 2017 spotlight paper. A third-party implementat

215 Dec 07, 2022

An Optical Character Recognition system using Pytesseract/Extracting data from Blood Pressure Reports.

Optical_Character_Recognition An Optical Character Recognition system using Pytesseract/Extracting data from Blood Pressure Reports. As an IOT/Compute

1 Feb 12, 2022

Detecting Text in Natural Image with Connectionist Text Proposal Network (ECCV'16)

Detecting Text in Natural Image with Connectionist Text Proposal Network The codes are used for implementing CTPN for scene text detection, described

1.3k Dec 22, 2022

A Vietnamese personal card OCR website built with Django.

Django VietCardOCR Installation Creation of virtual environments is done by executing the command venv: python -m venv venv That will create a new fol

4 Sep 04, 2021

A bot that extract text from images using the Tesseract OCR.

Text from image (OCR) @ocr_text_bot A simple bot to extract text from images. Usage What do I need? A AWS key configured locally, see here. NodeJS. I

4 Aug 06, 2021

YOLOv5 in DOTA with CSL_label.(Oriented Object Detection)（Rotation Detection）（Rotated BBox）

YOLOv5_DOTA_OBB YOLOv5 in DOTA_OBB dataset with CSL_label.(Oriented Object Detection) Datasets and pretrained checkpoint Datasets : DOTA Pretrained Ch

1.1k Dec 30, 2022

Code for AAAI 2021 paper: Sequential End-to-end Network for Efficient Person Search

This repository hosts the source code of our paper: [AAAI 2021]Sequential End-to-end Network for Efficient Person Search. SeqNet achieves the state-of

218 Dec 31, 2022

Maze generator and solver with python

Procedural-Maze-Generator-Algorithms Check out my youtube channel : Auctux Ressources Thanks to Jamis Buck Book : Mazes for programmers Requirements P

19 Dec 07, 2022

Unofficial implementation of "TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images"

TableNet Unofficial implementation of ICDAR 2019 paper : TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from

243 Dec 30, 2022

The project is an official implementation of our paper "3D Human Pose Estimation with Spatial and Temporal Transformers".

3D Human Pose Estimation with Spatial and Temporal Transformers This repo is the official implementation for 3D Human Pose Estimation with Spatial and

363 Dec 28, 2022

Fun program to overlay a mask to yourself using a webcam

Superhero Mask Overlay Description Simple project made for fun. It consists of placing a mask (a PNG image with transparent background) on your face.

10 Dec 01, 2022

Crop regions in napari manually

napari-crop Crop regions in napari manually Usage Create a new shapes layer to annotate the region you would like to crop: Use the rectangle tool to a

4 Sep 29, 2022

[BMVC'21] Official PyTorch Implementation of Grounded Situation Recognition with Transformers

Grounded Situation Recognition with Transformers Paper | Model Checkpoint This is the official PyTorch implementation of Grounded Situation Recognitio

18 Jul 19, 2022

Optical character recognition for Japanese text, with the main focus being Japanese manga

Manga OCR Optical character recognition for Japanese text, with the main focus being Japanese manga. It uses a custom end-to-end model built with Tran

327 Jan 01, 2023

Educational application aimed at automating user-defined workflows for the mobile game, "Granblue Fantasy", using a variety of CV technologies in the backend such as OpenCV, PyAutoGUI and EasyOCR and a frontend coded in Typescript.

Granblue Automation using Template Matching (It is like Full Auto, but with Full Customization!) Discord here: https://discord.gg/5Yv4kqjAbm Android v

71 Dec 30, 2022

AdvancedEAST is an algorithm used for Scene image text detect, which is primarily based on EAST, and the significant improvement was also made, which make long text predictions more accurate.https://github.com/huoyijie/raspberrypi-car

AdvancedEAST AdvancedEAST is an algorithm used for Scene image text detect, which is primarily based on EAST:An Efficient and Accurate Scene Text Dete

1.2k Dec 29, 2022

Python library to extract tabular data from images and scanned PDFs

Related tags

Overview

Overview

Prerequisite

Installation

Basic Usage

Detailed Library Usage

Woahh, as simple as that ?!

Explore

Bug Reports

License

Social Media

Comments

Releases(v2.4.0)

v2.4.0(Jul 18, 2022)

v2.3.1(May 6, 2022)

v2.2.0(Apr 20, 2021)

v2.1.2(Nov 6, 2020)

v2.1.0(Aug 27, 2020)

v2.0.2(Jul 4, 2020)

v2.0.1(Jul 3, 2020)

v2.0.0(Apr 30, 2020)

v1.2.1.2(Dec 1, 2019)

v1.1.0(Oct 20, 2019)

v1.0.1(Oct 7, 2019)

Owner

Org. Account

This tool will help you convert your text to handwriting xD

Single Shot Text Detector with Regional Attention

An Optical Character Recognition system using Pytesseract/Extracting data from Blood Pressure Reports.

Detecting Text in Natural Image with Connectionist Text Proposal Network (ECCV'16)

A Vietnamese personal card OCR website built with Django.

A bot that extract text from images using the Tesseract OCR.

YOLOv5 in DOTA with CSL_label.(Oriented Object Detection)（Rotation Detection）（Rotated BBox）

Code for AAAI 2021 paper: Sequential End-to-end Network for Efficient Person Search

Maze generator and solver with python

Unofficial implementation of "TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images"

The project is an official implementation of our paper "3D Human Pose Estimation with Spatial and Temporal Transformers".

Fun program to overlay a mask to yourself using a webcam

Crop regions in napari manually

[BMVC'21] Official PyTorch Implementation of Grounded Situation Recognition with Transformers

Optical character recognition for Japanese text, with the main focus being Japanese manga

Educational application aimed at automating user-defined workflows for the mobile game, "Granblue Fantasy", using a variety of CV technologies in the backend such as OpenCV, PyAutoGUI and EasyOCR and a frontend coded in Typescript.

Python tool that takes the OCR.space JSON output as input and draws a text overlay on top of the image.

virtual mouse which can copy files, close tabs and many other features !

A curated list of resources dedicated to scene text localization and recognition

AdvancedEAST is an algorithm used for Scene image text detect, which is primarily based on EAST, and the significant improvement was also made, which make long text predictions more accurate.https://github.com/huoyijie/raspberrypi-car