Skip to main content

License: MIT PyPI PyPI_versions PyPI_status PyPI_format Unit Tests Docs

VisArchPy

Data pipelines for extraction, transformation and visualization of architectural visuals in Python. It extracts images embedded in PDF files, collects relevant metadata, and extracts visual features using the DinoV2 model. We ambition to make of this package Ai-powered tool with features for recorgnizing different types architectural visuals (types of buildings, structures, etc.). The package is still in development and we are working on adding more features and improving the existing ones. If you have any suggestions or questions, please open an issue in our GitHub repository.

Main Features

Extraction pipelines

  • Layout: pipeline for extracting metadata and visuals (images) from PDF files using a layout analysis. Layout analysis recursively checks elements in the PDF file and sorts them into images, text, and other elements.
  • OCR: pipeline for extracting metadata and visuals from PDF files using OCR analysis. OCR analysis extracts images from PDF files using Tesseract OCR.
  • LayoutOCR: pipeline for extracting metadata and visuals from PDF files that combines layout and OCR analysis.

Metadata Extraction

  • Extraction of medatdata of extracted images (document page, image size)
  • Extraction of captions of images based on proximity to images and text-analysis using keywords.

Transformation utilities

  • Dino: pipeline for transforming images into visual features using the self-supervised learning in DinoV2.

Visualization utilities

  • Viz: an utility to create a bounding box plot. This plot provides an overview of the shapes and sizes of images in a data set.

    Example Bbox plot

Dependencies

Installion

After installing the dependencies, install VisArchPy using pip.

pip install visarchpy

Installing from source

  1. Clone the repository.

    git clone https://github.com/AiDAPT-A/VisArchPy.git
    
  2. Go to the root of the repository.

    cd VisArchPy/
    
  3. Install the package using pip.

    pip install .
    

Developers who intend to modify the sourcecode can install additional dependencies for test and documentation as follows.

  1. Go to the root directory visarchpy/

  2. Run:

pip install -e .[dev]

Usage

VisArchPy provides a command line interface to access its functionality. If you want to VisArchPy as a Python package consult the documentation.

  1. To access the CLI:
visarch -h
  1. To access a particular pipeline:
visarch [PIPELINE] [SUBCOMMAND]

For example, to run the layout pipeline using a single PDF file, do the following:

visarch layout from-file <path-to-pdf-file> <path-output-directory>

Use visarch [PIPELINE] [SUBCOMMAND] -h for help.

Results

Results from the data extraction pipelines (Layout, OCR, LayoutOCR) are save to the output directory. Results are organized as following:

00000/  # results directory
├── pdf-001  # directory where images are saved to. One per PDF file
├── 00000-metadata.csv  # extracted metadata as CSV
├── 00000-metadata.json  # extracted metadata as JSON
├── 00000-settings.json  # settings used by pipeline
└── 00000.log  # log file

Settings

The pipeline's settings determine how visual extraction from PDF files is performed. Settings must be passed as a JSON file on the CLI. Settings may must include all items listed below. The values showed belowed are the defaults.

Available settings
{
    "layout": { # setting for layout analysis
        "caption": { 
            "offset": [ # distance used to locate captions
                4,
                "mm"
            ],
            "direction": "down", # direction used to locate captions
            "keywords": [  # keywords used to find captions based on text analysis
                "figure",
                "caption",
                "figuur"
            ]
        },
        "image": { # images smaller than these dimensions will be ignored
            "width": 120,
            "height": 120
        }
    },
    "ocr": {  # settings for OCR analysis
        "caption": {
            "offset": [
                50,
                "px"
            ],
            "direction": "down",
            "keywords": [
                "figure",
                "caption",
                "figuur"
            ]
        },
        "image": {
            "width": 120,
            "height": 120
        },
        "resolution": 250, # dpi to convert PDF pages to images before OCR
        "resize": 30000  # total pixels. Larger OCR inputs are downsize to this before OCR
        "tesseract" : "--psm 1 --oem 3"  # tesseract options
    }
}


When no seetings are passed to a pipeline, the defaults are used. To print the default seetting to the terminal use:

visarch [PIPELINE] settings

Citation

Please cite this software using as follows:

Garcia Alvarez, M. G., Khademi, S., & Pohl, D. (2023). VisArchPy [Computer software]. https://github.com/AiDAPT-A/VisArchPy

Acknowlegdements

  • VisArchPy was develped thanks to the support provided by the Digital Competence Centre, Delft University of Technology.
  • Reseach Data Services, Delft University of Technology, The Netherlands.

Release files for visarchpy 1.0.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for visarchpy 1.0.4
File Size Uploaded
visarchpy-1.0.4.tar.gz 39.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for visarchpy 1.0.4
File Interpreter ABI Platform
visarchpy-1.0.4-py3-none-any.whl Python 3 none any Details

Total release size: 81.2 kB

Release files / visarchpy-1.0.4.tar.gz

Download URL visarchpy-1.0.4.tar.gz
Size 39.8 kB
Tags Source
SHA-256 checksum
How to use checksums
5ddb95e7d862d65659c22af4879b2eebbe32485c37d9c7ee6de580a98d16cbff
BLAKE2b-256 checksum
How to use checksums
f6d359d0edff4ac90338e90b0966462e6ea1639ba44cf9d8276d58ea20f1cd93
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.18

Release files / visarchpy-1.0.4-py3-none-any.whl

Download URL visarchpy-1.0.4-py3-none-any.whl
Size 41.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
31f6da75819e67f2c7da9dc43532884d569debfed28309734916f38ebb624056
BLAKE2b-256 checksum
How to use checksums
a93a2ba8695d78580de421ba65a209cc56464006b4dfadc903b72cc221da0bb3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.18

Release history Release notifications | RSS feed

This release

1.0.4 This release

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page