Doc-UFCN

Project description

Doc-UFCN

This Python 3 library contains a public implementation of Doc-UFCN, a fully convolutional network presented in the paper Multiple Document Datasets Pre-training Improves Text Line Detection With Deep Neural Networks. This library has been developed by the original authors from Teklia.

The model is designed to run various Document Layout Analysis (DLA) tasks like the text line detection or page segmentation.

Model schema

This library can be used by anyone that has an already trained Doc-UFCN model and want to easily apply it to document images. With only a few lines of code, the trained model is loaded, applied to an image and the detected objects along with some visualizations are obtained.

Getting started

To use Doc-UFCN in your own scripts, install it using pip:

pip install doc-ufcn

Usage

To apply Doc-UFCN to an image, one need to first add a few imports and to load an image. Note that the image should be in RGB.

import cv2
from doc_ufcn.main import DocUFCN

image = cv2.cvtColor(cv2.imread(IMAGE_PATH), cv2.COLOR_BGR2RGB)

Then one can initialize and load the trained model with the parameters used during training. The number of classes should include the background that must have been put as the first channel during training. By default, the model is loaded in evaluation mode. To load it in training mode, use mode="train".

nb_of_classes = 2
mean = [0, 0, 0]
std = [1, 1, 1]
input_size = 768
model_path = "trained_model.pth"

model = DocUFCN(nb_of_classes, input_size, 'cpu')
model.load(model_path, mean, std, mode="eval")

To run the inference on a GPU, one can replace cpu by the name of the GPU. In the end, one can run the prediction:

detected_polygons = model.predict(image)

Output

When running inference on an image, the detected objects are returned as in the following example. The objects belonging to a class (except for the background class) are returned as a list containing the confidence score and the polygon coordinates of each object.

{
  1: [
    {
      'confidence': 0.99,
      'polygon': [(490, 140), (490, 1596), (2866, 1598), (2870, 140)]
    }
    ...
  ],
  ...
}

In addition, one can directly retrieve the raw probabilities output by the model using model.predict(image, raw_output=True). A tensor of size (nb_of_classes, height, width) is then returned along with the polygons and can be used for further processing.

Lastly, two visualizations can be returned by the model:

A mask of the detected objects mask_output=True;
An overlap of the detected objects on the input image overlap_output=True.

By default, only the detected polygons are returned, to return the four outputs, one can use:

detected_polygons, probabilities, mask, overlap = model.predict(
    image, raw_output, mask_output, overlap_output
)

Mask of detected objects Overlap with the detected objects

Cite us!

If you want to cite us in one of your works, please use the following citation.

@inproceedings{boillet2020,
    author = {Boillet, Mélodie and Kermorvant, Christopher and Paquet, Thierry},
    title = {{Multiple Document Datasets Pre-training Improves Text Line Detection With
              Deep Neural Networks}},
    booktitle = {2020 25th International Conference on Pattern Recognition (ICPR)},
    year = {2021},
    month = Jan,
    pages = {2134-2141},
    doi = {10.1109/ICPR48806.2021.9412447}
}

License

This library is under the 3-Clause BSD License.

Project details

Release history Release notifications | RSS feed

0.2.0rc3 pre-release

Apr 12, 2024

0.2.0rc2 pre-release

Mar 1, 2024

0.2.0rc1 pre-release

Feb 28, 2024

0.1.9

Nov 13, 2023

0.1.9rc8 pre-release

Nov 13, 2023

0.1.9rc7 pre-release

Nov 7, 2023

0.1.9rc6 pre-release

Aug 21, 2023

0.1.9rc5 pre-release

Aug 21, 2023

0.1.9rc4 pre-release

Apr 13, 2023

0.1.9rc3 pre-release

Apr 13, 2023

0.1.9rc2 pre-release

Feb 24, 2023

0.1.9rc1 pre-release

Feb 16, 2023

0.1.8

Jan 18, 2023

0.1.8rc5 pre-release

Jan 16, 2023

0.1.8rc4 pre-release

Nov 30, 2022

0.1.8rc3 pre-release

Nov 29, 2022

0.1.8rc2 pre-release

Nov 14, 2022

0.1.8rc1 pre-release

Aug 26, 2022

0.1.7

Jul 4, 2022

0.1.5

Apr 1, 2022

0.1.4

Jan 26, 2022

0.1.3

Dec 1, 2021

This version

0.1.2

Nov 12, 2021

0.1.1

Nov 10, 2021

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

doc-ufcn-0.1.2.tar.gz (14.1 kB view hashes)

Uploaded Nov 12, 2021 Source

Built Distribution

doc_ufcn-0.1.2-py3-none-any.whl (15.2 kB view hashes)

Uploaded Nov 12, 2021 Python 3

Hashes for doc-ufcn-0.1.2.tar.gz

Hashes for doc-ufcn-0.1.2.tar.gz
Algorithm	Hash digest
SHA256	`eea124afb49cda24cb6e73db1f4acd7e309d0cee4b599f88a9f0b9d2ff0f75e1`
MD5	`d12379b879a37df7d1a034356b693063`
BLAKE2b-256	`5c4fc7714585df63204a468877d1ac87d482f56b4844ece214c19d32900e59b9`

Hashes for doc_ufcn-0.1.2-py3-none-any.whl

Hashes for doc_ufcn-0.1.2-py3-none-any.whl
Algorithm	Hash digest
SHA256	`cf59549d5212aba38990947cf48a31770ff28f9123f56220933ce3863e7fe3a2`
MD5	`f3f4f1146ebc35552db1cc000657026a`
BLAKE2b-256	`7dfd797ca4c78a91708456e57cf0025c43eb56074b709a73ef4175807d42820b`