Vision API — Python client
Official Python client for Vision API — send an image or a PDF, describe the fields you want in plain language, get structured JSON back with a confidence level on every value.
- Website — https://visionapi.io
- Documentation — https://docs.visionapi.io
- API keys — https://app.visionapi.io/dashboard/keys
- Preset catalog — https://visionapi.io/presets
- Playground — https://visionapi.io/playground
- Support — https://support.visionapi.io · https://visionapi.io/contact-us
Install
pip install visionapi-client
The distribution is visionapi-client; the module you import is visionapi.
Python 3.9+. Standard library only — no requests, no httpx, nothing to conflict with
what your project already pins.
Quick start
from visionapi import VisionAPI
vision = VisionAPI() # reads $VISION_API_KEY
res = vision.analyze(file="invoice.pdf", preset="invoice")
print(res["result"]["invoice_id"]["value"]) # 'A-10422'
print(res["result"]["total"]["value"]) # 1284.5 — or None, if the invoice has no total
print(res["credits_used"], res["credits_remaining"])
Requests are metered in credits, per image and per selected PDF page — see pricing for current rates. Failures cost nothing: the reservation is released in full on any non-2xx, so there is no compensating logic to write.
Server-side only. There is no publishable key and no test mode — an API key is a live spending credential. Never ship one to a browser, a mobile app or a notebook you share.
Reading a result
Responses are plain dictionaries, so everything you already know about dicts applies. Two rules explain almost every surprise:
1. Every scalar is wrapped. {"value": …, "confidence": "low"|"mid"|"high"}. Read
res["result"]["total"]["value"], not res["result"]["total"].
2. A preset response contains every field of that preset — including the ones the
document does not carry, which come back as {"value": None, "confidence": "low"}. A key
being present does not mean a value was found. Check value is not None.
Line-item arrays are the one shape worth looking at twice. The array itself is not wrapped; each cell inside each row is:
{
"invoice_id": {"value": "A-10422", "confidence": "high"},
"carrier": {"value": None, "confidence": "low"},
"line_item": [
{"description": {"value": "Widget", "confidence": "high"},
"quantity": {"value": 2, "confidence": "high"},
"amount": {"value": 25.0, "confidence": "mid"}},
],
}
Helpers ship for the common readings, so you rarely have to spell that out:
from visionapi import unwrap, value, rows, present, missing, below_confidence
unwrap(res["result"])
# {'invoice_id': 'A-10422', 'carrier': None, 'line_item': [{'description': 'Widget', …}]}
unwrap(res["result"], drop_null=True) # only what was actually found
value(res["result"], "total", 0) # 1284.5, or 0 when absent
rows(res["result"], "line_item") # [] when the invoice has no lines
present(res["result"]) # ['invoice_id', 'total', 'line_item']
missing(res["result"]) # ['carrier', …]
below_confidence(res["result"], "high") # fields to route to a human
TypedDict definitions for every response live in visionapi.types, so mypy and your
editor know the shape without turning responses into objects you have to unwrap twice.
What you can send
Exactly one file source per call:
vision.analyze(file="invoice.pdf", preset="invoice") # a path
vision.analyze(file=open("invoice.pdf", "rb"), preset="invoice") # an open binary file
vision.analyze(file=raw_bytes, preset="invoice") # bytes
vision.analyze(file=("scan.png", raw_bytes), preset="invoice") # bytes + a name
vision.analyze(file_url="https://example.com/invoice.pdf", preset="invoice")
vision.analyze(file_base64=b64, preset="invoice") # `data:` prefix optional
JPEG, PNG, WebP, TIFF and PDF, up to 20 MB and 50 pages. The type is detected from magic bytes — the filename is ignored.
Options
| Argument | Default | What it does |
|---|---|---|
preset |
— | A catalog name, or "auto" to let the API classify the file first (free). |
schema |
— | Custom fields, alone or on top of a preset. |
schema_name |
— | A schema saved in your dashboard. Excludes preset and schema. |
pages |
all | PDF page selection, e.g. "1-3,7". You pay for selected pages only. |
language_hint |
auto | ISO 639-1 code, e.g. "es". |
detail |
"standard" |
"high" renders pages at higher resolution. Same cost, slower. |
output |
"json" |
"text" returns raw OCR text instead of fields. |
include_raw_text |
False |
Adds full_text, the whole transcription, alongside result. |
min_confidence |
"low" |
Fields below the level come back None, with confidence preserved. |
Custom fields
A schema is a flat dict: each key is a field name, each value describes what to extract. It is compiled before any credit moves, so a bad schema costs nothing.
res = vision.analyze(
file="invoice.pdf",
preset="invoice",
schema={
# Plain form — the string is the description, type defaults to string.
"machine_serial": 'Serial number of the machine being invoiced, without the "SN:" prefix',
# Typed form.
"total_net": {"type": "number", "description": "Total before tax"},
"signed_on": {"type": "date", "description": "Date the contract was signed"},
"is_paid": {"type": "boolean", "description": "Whether the invoice is stamped PAID"},
# Reserved key: injects fields into every row of the preset's line-item array.
"line_item": {"lot_number": "The lot number printed on the line, if present"},
},
)
Field names must match ^[a-z][a-z0-9_]{0,63}$. Types are string (default), number,
boolean, date, array and object. A custom name that collides with a preset field is
a 422 schema_field_conflict — rename it, or use the preset's own field.
Descriptions are the prompt. "The invoice number exactly as printed, without the #"
extracts better than "invoice number". Say what to do when the value is missing or
ambiguous if it matters.
Reuse a combination by saving it:
vision.create_schema("our-invoices", preset="invoice", schema={"machine_serial": "…"})
vision.analyze(file="invoice.pdf", schema_name="our-invoices")
Picking a preset
28 presets ship with the API. Fetch the catalog rather than hardcoding field names from memory — presets are versioned, and the catalog is the source of truth:
for p in vision.presets(): # no API key required
print(p["name"], p["kind"], p["field_count"])
invoice = vision.preset("invoice")
[f["name"] for f in invoice["fields"]]
Three ways to choose:
# 1. You know what it is.
vision.analyze(file="receipt.jpg", preset="receipt")
# 2. You don't, and you want the data anyway. Classification is free.
res = vision.analyze(file="unknown.pdf", preset="auto")
res["detection"]["preset"] # what ran
res["detection"]["fallback"] # True = "shape unknown", not a match
res["detection"]["alternatives"] # the rest of the ranking, best first
# 3. The *type* is the decision — routing a mixed inbox, or refusing to spend
# on a 40-page PDF until you know what it is. Far cheaper than extracting.
guess = vision.detect(file="unknown.pdf")
if guess["recommended"] == "invoice" and not guess["fallback"]:
vision.analyze(file="unknown.pdf", preset="invoice")
detect reads page 1 only, so an image and a 300-page PDF cost the same, and it is metered
in batches rather than per call: most calls report credits_used: 0 and an occasional one
carries the charge. See pricing for the rate.
Questions instead of fields
Up to 5 questions about one file, priced exactly like an extraction. The questions themselves are free.
res = vision.ask(
file="photo.jpg",
questions=["Is there a dog in the image?", "How many people are visible?"],
)
for a in res["answers"]:
print(a["question"], "→", a["verdict"], a["answer"])
verdict is "yes", "no", "uncertain" (a yes/no question the image does not settle)
or "n/a" (not a yes/no question). Branch on it instead of parsing the prose.
Long jobs: async and webhooks
Synchronous requests are killed at 60 seconds with a 504 sync_timeout. Anything that
might run longer — a long PDF, detail="high", a batch — belongs on the queue.
# Submit, then poll. wait_for_task handles the loop and the failure case.
task = vision.analyze_and_wait(
file="contract-80-pages.pdf",
preset="contract",
pages="1-50",
poll_interval=2.0,
max_wait=900,
on_poll=lambda t: print(t["status"]),
)
print(task["result"]["parties"]["value"])
# Or submit and walk away — the result comes to you.
ref = vision.analyze_async(
file="contract.pdf",
preset="contract",
webhook_url="https://yourapp.com/hooks/vision",
)
Results stay retrievable for 7 days; after that the task raises ResultExpiredError
(metadata survives, the payload does not).
Verifying a delivery
Deliveries are signed. Verify over the raw bytes before parsing — a re-serialized body has different bytes and will not match.
from flask import Flask, request
from visionapi import verify_webhook, WebhookSignatureError
app = Flask(__name__)
SECRET = os.environ["VISION_WEBHOOK_SECRET"]
@app.post("/hooks/vision")
def hook():
try:
event = verify_webhook(request.get_data(), request.headers.get("X-Vision-Signature"), SECRET)
except WebhookSignatureError:
return "", 400 # never parse an unverified body
queue.put(event) # event["event"] is 'task.completed' | 'task.failed'
return "", 202 # any 2xx is success — ack fast, work afterwards
verify_webhook rejects a bad signature, a malformed header and a timestamp more than 5
minutes old, and accepts a delivery if any v1= part matches — which is what makes a
secret rotation seamless. Get the secret from
https://app.visionapi.io/dashboard/webhooks. Failed deliveries retry at +1 m, +5 m,
+15 m and +40 m, then stop.
Errors
Every failure raises a subclass of VisionAPIError carrying the HTTP status, the stable
code, and whatever details the endpoint attached. Branch on the class or on code —
never on the message text, which is prose and changes.
from visionapi import (
InsufficientCreditsError,
RateLimitError,
SyncTimeoutError,
UnsupportedTypeError,
VisionAPIError,
)
try:
res = vision.analyze(file="scan.pdf", preset="invoice")
except InsufficientCreditsError as e:
alert_ops(f"needs {e.required}, has {e.available}") # never retried — it cannot succeed
except SyncTimeoutError:
task = vision.analyze_and_wait(file="scan.pdf", preset="invoice")
except UnsupportedTypeError:
quarantine("not an image or a PDF")
except VisionAPIError as e:
log.error("vision failed", code=e.code, status=e.status, request_id=e.request_id)
| Exception | HTTP | Codes |
|---|---|---|
InvalidRequestError |
400 | invalid_request |
AuthenticationError |
401 | invalid_api_key, unauthorized |
InsufficientCreditsError |
402 | insufficient_credits — with .required / .available |
PermissionDeniedError |
403 | forbidden, email_not_verified |
NotFoundError |
404 | task_not_found, schema_not_found |
ConflictError |
409 | conflict |
ResultExpiredError |
410 | result_expired |
PayloadTooLargeError |
413 | file_too_large, page_limit_exceeded |
UnsupportedTypeError |
415 | unsupported_type |
UnprocessableError |
422 | pdf_encrypted, invalid_page_selection, invalid_schema, schema_field_conflict, too_many_questions |
RateLimitError |
429 | rate_limited — with .retry_after |
TooManyTasksError |
429 | too_many_tasks — the per-plan async concurrency cap, with .max_tasks. Subclasses RateLimitError, but is not auto-retried: it clears when one of your tasks finishes |
InternalError |
500 | internal_error — with .request_id |
ProviderError |
502 | provider_error |
SyncTimeoutError |
504 | sync_timeout |
UsageError (bad arguments), APIConnectionError / APITimeoutError (the request never
got a response) and TaskFailedError / TaskTimeoutError come from the client itself.
They are named that way deliberately — shadowing the builtin ConnectionError,
TimeoutError and PermissionError in your except clauses would be a nasty surprise.
Retries and idempotency
The client retries 429, 500, 502 and network failures — max_retries=3 by default, with
the server's own Retry-After honored on 429 and exponential backoff with jitter
elsewhere. Input errors and insufficient_credits are never retried, because they cannot
succeed.
Every billable POST is sent with a generated Idempotency-Key, so a retried upload replays
the first response instead of paying twice. Supply your own when the caller may retry — a
job that re-runs, a queue that redelivers — because a fresh process generates a fresh key:
vision.analyze(file=path, preset="invoice", idempotency_key=f"invoice-{invoice_id}")
Reusing a key with a different payload raises ConflictError, which is the mechanism
working: it means the key already stands for something else.
Configuration
vision = VisionAPI(
api_key=os.environ["VISION_API_KEY"], # default: $VISION_API_KEY
base_url="https://api.visionapi.io", # default; override for a self-hosted deployment
timeout=120.0, # per request, seconds
max_retries=3,
auto_idempotency=True,
headers={"x-trace-id": trace_id}, # sent on every request
)
Every method takes per-call idempotency_key= and timeout=.
TLS certificates. The client uses certifi's CA
bundle when it is importable, and the system store otherwise. That is deliberate: a
python.org macOS build ships with an empty store until you run
Install Certificates.command, and the resulting SSLCertVerificationError looks like an
API problem rather than an interpreter one. pip install certifi fixes it; set
$VISION_CA_BUNDLE to point at your own root if you are behind a TLS-inspecting proxy.
Account and usage
credits = vision.credits()
credits["balance"], credits["buckets"]
# buckets are spent in order: subscription → rollover → pack → welcome
for record in vision.iter_requests(limit=100):
print(record["created_at"], record["endpoint"], record["preset"], record["credits_used"])
Usage history is metadata only — never the file, never the extracted values. Uploaded files are never retained: a synchronous request holds yours in memory for the length of the call, and an async request stages it only until the worker finishes with it.
Limits
Same for everyone:
| Limit | Value |
|---|---|
| Max file size | 20 MB |
| Max PDF pages per request | 50 |
| Sync request timeout | 60 s |
Per plan:
| Limit | Free | Starter | Growth | Pro | Scale |
|---|---|---|---|---|---|
| Requests per minute, per key | 10 | 60 | 120 | 300 | 600 |
| Burst capacity | 20 | 120 | 240 | 600 | 1,200 |
| Concurrent async tasks | 1 | 4 | 8 | 16 | 32 |
| Active API keys per account | 1 | 5 | 10 | 20 | 50 |
| Saved schemas | 3 | 10 | 25 | 100 | unlimited |
Max questions per ask |
5 | 5 | 5 | 10 | 10 |
The rate-limit bucket is per API key, not per account — splitting a workload across
keys splits the limit too. The concurrency cap is per account and does not split that way:
over it, an async submission answers 429 too_many_tasks and is charged nothing.
Higher limits on paid plans: https://visionapi.io/pricing.
Examples
Runnable scripts in examples/:
| File | What it shows |
|---|---|
analyze.py |
The smallest useful call, and how to read the result |
custom_schema.py |
Custom fields, line-item injection, saved schemas |
detect_then_analyze.py |
Routing a mixed inbox before spending on extraction |
async_batch.py |
A folder of long PDFs, queued with bounded concurrency |
webhook_server.py |
A verified receiver, with no framework |
ask.py |
Visual Q&A and the verdict field |
dataframe.py |
Line items → pandas DataFrame → CSV |
export VISION_API_KEY=sk_live_…
python examples/analyze.py invoice.pdf
Development
pip install -e ".[dev]"
pytest # offline: the transport is stubbed, no key and no network needed
mypy src
ruff check .
Contributing
Issues and pull requests are welcome at https://github.com/devrobotlabs/visionapi-python. For anything about the API itself — a preset, a limit, an error code — https://support.visionapi.io reaches the team faster.
License
MIT © Vision API
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file visionapi_client-1.0.0.tar.gz.
File metadata
- Download URL: visionapi_client-1.0.0.tar.gz
- Upload date:
- Size: 37.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1079b44fbfb495402187be06f51791e7c1c9feeb3873a99d9a245d7c43cf0f4f
|
|
| MD5 |
e438dbfb5301db3860159ce9636b305e
|
|
| BLAKE2b-256 |
2dcb5055695c04e65dd45a998f3cbd14aa938209842ed59003616a9127f78a76
|
Provenance
The following attestation bundles were made for visionapi_client-1.0.0.tar.gz:
Publisher:
publish.yml on devrobotlabs/visionapi-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
visionapi_client-1.0.0.tar.gz -
Subject digest:
1079b44fbfb495402187be06f51791e7c1c9feeb3873a99d9a245d7c43cf0f4f - Sigstore transparency entry: 2429635805
- Sigstore integration time:
-
Permalink:
devrobotlabs/visionapi-python@14d34a5244a504632342e6a7b2dec44f8bf333ba -
Branch / Tag:
refs/tags/v1.0.0 - Owner: https://github.com/devrobotlabs
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@14d34a5244a504632342e6a7b2dec44f8bf333ba -
Trigger Event:
push
-
Statement type:
File details
Details for the file visionapi_client-1.0.0-py3-none-any.whl.
File metadata
- Download URL: visionapi_client-1.0.0-py3-none-any.whl
- Upload date:
- Size: 31.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ad2205603227031c6c8f8c36d3d0fde363def3949f373fd390c6f59e14995c6e
|
|
| MD5 |
44a231febcf2a3827730578eb044b814
|
|
| BLAKE2b-256 |
81c523eada0b52a6d5e6611a3c103e95878c8c42dfa67857b1aacd0bcb0e3182
|
Provenance
The following attestation bundles were made for visionapi_client-1.0.0-py3-none-any.whl:
Publisher:
publish.yml on devrobotlabs/visionapi-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
visionapi_client-1.0.0-py3-none-any.whl -
Subject digest:
ad2205603227031c6c8f8c36d3d0fde363def3949f373fd390c6f59e14995c6e - Sigstore transparency entry: 2429635868
- Sigstore integration time:
-
Permalink:
devrobotlabs/visionapi-python@14d34a5244a504632342e6a7b2dec44f8bf333ba -
Branch / Tag:
refs/tags/v1.0.0 - Owner: https://github.com/devrobotlabs
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@14d34a5244a504632342e6a7b2dec44f8bf333ba -
Trigger Event:
push
-
Statement type: