invoice2data

Python parser to extract data from pdf invoice

Project description

# Data extractor for PDF invoices - invoice2data

[![Circle CI](https://circleci.com/gh/m3nu/invoice2data.svg?style=svg)](https://circleci.com/gh/m3nu/invoice2data)

A Python library to support your accounting process. Tested on Python 2.7, 3.4 and 3.5

- extracts text from PDF files
- searches for regex in the result
- saves results as CSV
- optionally renames PDF files to match the content

With the flexible template system you can:

- precisely match PDF files
- define static fields that are the same for every invoice
- have multiple regex per field (if layout or wording changes)
- define currency

Go from PDF files to this:

```
{'date': (2014, 5, 7), 'invoice_number': '30064443', 'amount': 34.73, 'desc': 'Invoice 30064443 from QualityHosting', 'lines': [{'price': 42.0, 'desc': u'Small Business StandardExchange 2010\nGrundgeb\xfchr pro Einheit\nDienst: OUDJQ_office\n01.05.14-31.05.14\n', 'pos': u'7', 'qty': 1.0}]}
{'date': (2014, 6, 4), 'invoice_number': 'EUVINS1-OF5-DE-120725895', 'amount': 35.24, 'desc': 'Invoice EUVINS1-OF5-DE-120725895 from Amazon EU'}
{'date': (2014, 8, 3), 'invoice_number': '42183017', 'amount': 4.11, 'desc': 'Invoice 42183017 from Amazon Web Services'}
{'date': (2015, 1, 28), 'invoice_number': '12429647', 'amount': 101.0, 'desc': 'Invoice 12429647 from Envato'}
```

## Installation

1. Install pdftotext

If possible get the latest [xpdf/poppler-utils](https://poppler.freedesktop.org/) version. It's included with OSX Homebrew, Debian Sid and Ubuntu 16.04. Without it, `pdftotext` won't parse tables in PDF correctly.

2. Install `invoice2data` using pip

```
pip install invoice2data
```

Optionally this uses `pdfminer`, but `pdftotext` works better. You can choose which module to use. No special Python packages are necessary at the moment, except for `pdftotext`.

There is also `tesseract` integration as a fallback, if no text can be extracted. But it may be more reliable to use

## Usage

Basic usage. Process PDF files and write result to CSV.
- `invoice2data invoice.pdf`
- `invoice2data *.pdf`

Specify folder with yml templates. (e.g. your suppliers)
`invoice2data --template-folder ACME-templates invoice.pdf`

Only use your own templates and exclude built-ins
`invoice2data --exclude-built-in-templates --template-folder ACME-templates invoice.pdf`

Processes a folder of invoices and copies renamed invoices to new folder.
`invoice2data --copy new_folder folder_with_invoices/*.pdf`

Processes a single file and dumps whole file for debugging (useful when adding new templates in templates.py)
`invoice2data --debug my_invoice.pdf`

Recognize test invoices:
`invoice2data invoice2data/test/pdfs/* --debug`

If you want to use it as a lib just do

```
from invoice2data import extract_data

result = extract_data('path/to/my/file.pdf')
```

## Template system

See `invoice2data/templates` for existing templates. Just extend the list to add your own. If deployed by a bigger organisation, there should be an interface to edit templates for new suppliers. 80-20 rule. For a short tutorial on how to add new templates, see [TUTORIAL.md](TUTORIAL.md).

Templates are based on Yaml. They define one or more keywords to find the right template and regexp for fields to be extracted. They could also be a static value, like the full company name.

Template files are tried in alphabetical order.

We may extend them to feature options to be used during invoice processing.

Example:

```
issuer: Amazon Web Services, Inc.
keywords:
- Amazon Web Services
fields:
amount: TOTAL AMOUNT DUE ON.*\$(\d+\.\d+)
amount_untaxed: TOTAL AMOUNT DUE ON.*\$(\d+\.\d+)
date: Invoice Date:\s+([a-zA-Z]+ \d+ , \d+)
invoice_number: Invoice Number:\s+(\d+)
partner_name: (Amazon Web Services, Inc\.)
options:
remove_whitespace: false
currency: HKD
date_formats:
- '%d/%m/%Y'
lines:
start: Detail
end: \* May include estimated US sales tax
first_line: ^ (?P<description>\w+.*)\$(?P<price_unit>\d+\.\d+)
line: (.*)\$(\d+\.\d+)
last_line: VAT \*\*
```

## Roadmap and open tasks

- tutorial and documentation for template options.
- integrate with online OCR?
- try to 'guess' parameters for new invoice formats.
- can apply machine learning to guess new parameters?

## Maintainers
- [Manuel Riel](https://github.com/m3nu)
- [Alexis de Lattre](https://github.com/alexis-via): Add setup.py for Pypi, fix locale bug, add templates for new invoice types.

## Other Contributors
- [Holger Brunn](https://github.com/hbrunn): Add support for parsing invoice items.

Project details

Release history Release notifications | RSS feed

0.4.5

Nov 26, 2023

0.4.4

Apr 8, 2023

0.4.3

Mar 31, 2023

0.4.2

Feb 11, 2023

0.4.1

Feb 6, 2023

0.4.0

Dec 12, 2022

0.3.6

Jun 21, 2021

0.3.5

Aug 21, 2019

0.3.4

Jun 21, 2019

0.3.3

Jan 30, 2019

0.3.2

Nov 28, 2018

0.3.1

Nov 9, 2018

0.2.103

Nov 9, 2018

0.2.101

Sep 8, 2018

0.2.100

Aug 17, 2018

0.2.99

Aug 9, 2018

0.2.98

Jun 8, 2018

0.2.97

May 27, 2018

0.2.96

May 27, 2018

0.2.95

May 27, 2018

0.2.94

May 27, 2018

0.2.93

May 24, 2018

0.2.92

May 22, 2018

0.2.91

May 21, 2018

0.2.90

May 21, 2018

0.2.89

May 20, 2018

0.2.88

May 15, 2018

0.2.87

May 15, 2018

0.2.86

May 14, 2018

0.2.85

May 13, 2018

0.2.84

May 6, 2018

0.2.83

May 2, 2018

0.2.82

Apr 20, 2018

0.2.81

Mar 20, 2018

0.2.80

Mar 19, 2018

0.2.79

Mar 18, 2018

0.2.78

Mar 15, 2018

0.2.77

Mar 14, 2018

0.2.76

Feb 26, 2018

0.2.75

Feb 26, 2018

0.2.74

Feb 17, 2018

0.2.73

Feb 16, 2018

This version

0.2.72

Feb 15, 2018

0.2.71

Feb 15, 2018

0.2.70

Jan 23, 2018

0.2.69

Jan 10, 2018

0.2.67

Dec 1, 2017

0.2.66

Nov 7, 2017

0.2.65

Oct 3, 2017

0.2.64

Sep 29, 2017

0.2.63

Sep 29, 2017

0.2.62

Sep 26, 2017

0.2.61

Aug 31, 2017

0.2.59

Jul 4, 2017

0.2.58

Jun 20, 2017

0.2.56

Jun 14, 2017

0.2.55

May 31, 2017

0.2.54

May 24, 2017

0.2.53

May 18, 2017

0.2.51

Mar 29, 2017

0.2.49

Mar 23, 2017

0.2.47

Mar 8, 2017

0.2.45

Mar 8, 2017

0.2.44

Mar 8, 2017

0.2.43

Feb 3, 2017

0.2.42

Jan 23, 2017

0.2.41

Jan 4, 2017

0.2.40

Dec 29, 2016

0.2.39

Dec 16, 2016

0.2.38

Nov 13, 2016

0.2.36

Oct 6, 2016

0.2.34

Oct 4, 2016

0.2.33

Sep 30, 2016

0.2.31

Sep 30, 2016

0.2.30

Sep 28, 2016

0.2.29

Jun 25, 2016

0.2.28

Jun 7, 2016

0.2.27

May 25, 2016

0.2.26

May 14, 2016

0.2.25

May 14, 2016

0.2.24

May 14, 2016

0.2.22

May 14, 2016

0.2.21

May 14, 2016

0.2.20

May 14, 2016

0.2.19

May 14, 2016

0.2.18

May 14, 2016

0.2.17

May 14, 2016

0.2.16

May 14, 2016

0.2.15

May 14, 2016

0.2.14

Apr 3, 2016

0.2.13

Apr 2, 2016

0.2.10

Apr 2, 2016

0.2.9

Apr 2, 2016

0.2.8

Apr 2, 2016

0.2.5

Mar 30, 2016

0.2.4

Mar 30, 2016

0.2.3

Mar 30, 2016

0.2.2

Mar 30, 2016

0.2.1

Mar 30, 2016

0.2.0

Jan 23, 2016

0.1.2

Jan 2, 2016

0.0.1

Dec 26, 2015

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

invoice2data-0.2.72.tar.gz (368.1 kB view details)

Uploaded Feb 15, 2018 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

invoice2data-0.2.72-py2.7.egg (408.7 kB view details)

Uploaded Feb 15, 2018 Egg

File details

Details for the file invoice2data-0.2.72.tar.gz.

File metadata

Download URL: invoice2data-0.2.72.tar.gz
Upload date: Feb 15, 2018
Size: 368.1 kB
Tags: Source
Uploaded using Trusted Publishing? No

File hashes

Hashes for invoice2data-0.2.72.tar.gz
Algorithm	Hash digest
SHA256	`0df2a33d5f9355c69793aa064a150eb6f93b1ca0a21075d6b5f59e63634681ca`
MD5	`49d477eaaafaf9593210d09cdaaf5717`
BLAKE2b-256	`9754faefd8c3f1aafc6b20413812f0cdbfb200dfd27a12b79cf3a8d730dcf745`

See more details on using hashes here.

File details

Details for the file invoice2data-0.2.72-py2.7.egg.

File metadata

Download URL: invoice2data-0.2.72-py2.7.egg
Upload date: Feb 15, 2018
Size: 408.7 kB
Tags: Egg
Uploaded using Trusted Publishing? No

File hashes

Hashes for invoice2data-0.2.72-py2.7.egg
Algorithm	Hash digest
SHA256	`d80a86085d806771de3718b311a2b8341be29c4396d6112c39053066ecf7725f`
MD5	`83019dbbf9f13830704ac50e4a5da49b`
BLAKE2b-256	`e06e03e212a728bf164c17d40de45560e5f4cde56ae5296d13b9928bf0cbd9fe`

See more details on using hashes here.

invoice2data 0.2.72

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Project description

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes