invoicedataextraction-sdk
Official Python SDK for Invoice Data Extraction. Uploads your files, submits the extraction, waits for it to finish and hands you the rows as data, or a spreadsheet, in a few lines of code.
- Python 3.9 or later
Install
pip install invoicedataextraction-sdk
Quick Start
import json
import os
import sys
from invoicedataextraction import InvoiceDataExtraction
from invoicedataextraction.errors import SdkError, ApiResponseError
try:
client = InvoiceDataExtraction(
api_key=os.environ.get("INVOICE_DATA_EXTRACTION_API_KEY"),
)
result = client.extract(
folder_path="./invoices",
prompt="Extract invoice number, date, vendor name, and total amount",
output_structure="per_invoice",
json_typed_values=True,
console_output=True, # remove to disable console logging
)
if result["status"] == "completed":
for row in client.iterate_results(extraction_id=result["extraction_id"]):
print(row) # one dict per extracted row, keyed by your output columns
except (SdkError, ApiResponseError) as error:
print(json.dumps(error.body, indent=2), file=sys.stderr)
raise SystemExit(1)
extract(...) uploads your files (pass a folder_path or a list of files), submits the extraction, waits until it finishes and returns the result: the final status response from the API, for a completed, failed or cancelled extraction. iterate_results(...) then reads the extracted rows straight from the API, with amounts as numbers and empty cells as None because the extraction was submitted with json_typed_values; get_results(...) reads one page when you want to manage paging yourself. Check result["pages"]["failed_count"] to verify that all uploaded pages were processed, and result["review_needed"]["count"] for rows that need a human's check before you rely on the data.
To get a spreadsheet, add download={"formats": ["xlsx"], "output_path": "./output"} to the call and the file is saved when the extraction completes, or call download_output(...) later.
Generate an API key from your dashboard. Every account includes 50 free pages per month. Additional credits can be purchased on a pay-as-you-go basis with no subscription needed.
Staged Workflow
If you need control over individual steps, for example uploading files in one part of your system and extracting in another, use the lower-level methods:
import json
import os
import sys
from invoicedataextraction import InvoiceDataExtraction
from invoicedataextraction.errors import SdkError, ApiResponseError
try:
client = InvoiceDataExtraction(
api_key=os.environ.get("INVOICE_DATA_EXTRACTION_API_KEY"),
)
upload = client.upload_files(
files=["./invoice1.pdf", "./invoice2.pdf"],
console_output=True,
)
submitted = client.submit_extraction(
upload_session_id=upload["upload_session_id"],
file_ids=upload["file_ids"],
prompt="Extract invoice number and total",
output_structure="per_invoice",
json_typed_values=True,
)
result = client.wait_for_extraction_to_finish(
extraction_id=submitted["extraction_id"],
console_output=True,
)
for row in client.iterate_results(extraction_id=submitted["extraction_id"]):
print(row)
except (SdkError, ApiResponseError) as error:
print(json.dumps(error.body, indent=2), file=sys.stderr)
raise SystemExit(1)
Options, and stopping a run
Beside json_typed_values, extract(...) and submit_extraction(...) take output_language, review_needed_fill_color, affected_field_fill_color and send_completion_email, each applying to that extraction only. Without json_typed_values, every value in the JSON output and in the rows is a string.
Set ask_questions=True and the extraction can stop to ask when the documents leave something unsettled, instead of deciding on its own. Pass on_questions and the SDK calls it as on_questions(questions, status), sends back the answers it returns and carries on to the result; without it, extract(...) returns status: "input_required" with the questions for you to answer with answer_questions(...). The questions, the answer forms and the deadline are in the API reference.
def answer(questions, status):
return [
{"question_id": question["question_id"], "accept_recommended": True}
# or {"question_id": ..., "choice_id": "b"}, or {"question_id": ..., "text": "DD/MM/YYYY"}
for question in questions
]
result = client.extract(
folder_path="./invoices",
prompt="Extract invoice number, date, vendor name, and total amount",
output_structure="per_invoice",
json_typed_values=True,
ask_questions=True,
on_questions=answer,
)
# Stop an extraction that is still queued or processing
client.cancel_extraction(extraction_id=extraction_id)
Listing past extractions
# One-page browse with filters
page = client.list_extractions(
status="completed",
limit=50,
)
# Auto-paginating iterator over every matching extraction
for extraction in client.iterate_extractions(status="completed"):
print(extraction["extraction_id"], extraction["task_name"])
# Full record (the original prompt, options, full pages, full failure error, etc.)
result = client.get_extraction(extraction_id="...")
extraction = result["extraction"]
On listing methods, team admins can pass scope="team" to see extractions submitted by any team member; pass scope="own" to force own-only results.
Error Handling
SDK methods raise SdkError or ApiResponseError on failure. The structured error body is on error.body, with fields error.body["error"]["code"], error.body["error"]["message"], error.body["error"]["retryable"], and error.body["error"]["details"].
When an extraction task itself reaches a terminal state, extract(...) returns that response rather than raising: check result["status"] for "completed", "failed", or "cancelled" for tasks stopped from the web app or with cancel_extraction(...). See the full docs for details.
Documentation
- Python SDK docs: full method reference, parameters, return shapes, and examples
- REST API docs: endpoint-level documentation for direct HTTP integration
- Dashboard: manage API keys and view extraction results
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file invoicedataextraction_sdk-0.6.0.tar.gz.
File metadata
- Download URL: invoicedataextraction_sdk-0.6.0.tar.gz
- Upload date:
- Size: 23.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
016207ebe46a0e22094341199cd016ad61c2a054f6813d53187cbd62a0634a97
|
|
| MD5 |
e2c49a45130f42fff859f86da4315739
|
|
| BLAKE2b-256 |
e7b0bb8e37aa2676fe7fcf72dcc1d0b1540a7f0d62626832e6a8e183e11b2582
|
File details
Details for the file invoicedataextraction_sdk-0.6.0-py3-none-any.whl.
File metadata
- Download URL: invoicedataextraction_sdk-0.6.0-py3-none-any.whl
- Upload date:
- Size: 23.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f09c0479bc830315af085d127b728c5ec33a9baff6e2eb1aedc382a89f120280
|
|
| MD5 |
77c28978b86b2ce8dc3748f120ecb8a6
|
|
| BLAKE2b-256 |
f8944ca549bcfd9ed7bd262a1fc03ba305327849fb59a30aa135d0ccf7e58205
|