Skip to main content

Vision API — Python client

Official Python client for Vision API — send an image or a PDF, describe the fields you want in plain language, get structured JSON back with a confidence level on every value.

PyPI Python versions license


Install

pip install visionapi-client

The distribution is visionapi-client; the module you import is visionapi.

Python 3.9+. Standard library only — no requests, no httpx, nothing to conflict with what your project already pins.

Quick start

from visionapi import VisionAPI

vision = VisionAPI()  # reads $VISION_API_KEY

res = vision.analyze(file="invoice.pdf", preset="invoice")

print(res["result"]["invoice_id"]["value"])  # 'A-10422'
print(res["result"]["total"]["value"])       # 1284.5 — or None, if the invoice has no total
print(res["credits_used"], res["credits_remaining"])

Requests are metered in credits, per image and per selected PDF page — see pricing for current rates. Failures cost nothing: the reservation is released in full on any non-2xx, so there is no compensating logic to write.

Server-side only. There is no publishable key and no test mode — an API key is a live spending credential. Never ship one to a browser, a mobile app or a notebook you share.


Reading a result

Responses are plain dictionaries, so everything you already know about dicts applies. Two rules explain almost every surprise:

1. Every scalar is wrapped. {"value": …, "confidence": "low"|"mid"|"high"}. Read res["result"]["total"]["value"], not res["result"]["total"].

2. A preset response contains every field of that preset — including the ones the document does not carry, which come back as {"value": None, "confidence": "low"}. A key being present does not mean a value was found. Check value is not None.

Line-item arrays are the one shape worth looking at twice. The array itself is not wrapped; each cell inside each row is:

{
  "invoice_id": {"value": "A-10422", "confidence": "high"},
  "carrier":    {"value": None,      "confidence": "low"},
  "line_item": [
    {"description": {"value": "Widget", "confidence": "high"},
     "quantity":    {"value": 2,        "confidence": "high"},
     "amount":      {"value": 25.0,     "confidence": "mid"}},
  ],
}

Helpers ship for the common readings, so you rarely have to spell that out:

from visionapi import unwrap, value, rows, present, missing, below_confidence

unwrap(res["result"])
# {'invoice_id': 'A-10422', 'carrier': None, 'line_item': [{'description': 'Widget', …}]}

unwrap(res["result"], drop_null=True)   # only what was actually found
value(res["result"], "total", 0)        # 1284.5, or 0 when absent
rows(res["result"], "line_item")        # [] when the invoice has no lines
present(res["result"])                  # ['invoice_id', 'total', 'line_item']
missing(res["result"])                  # ['carrier', …]
below_confidence(res["result"], "high") # fields to route to a human

TypedDict definitions for every response live in visionapi.types, so mypy and your editor know the shape without turning responses into objects you have to unwrap twice.


What you can send

Exactly one file source per call:

vision.analyze(file="invoice.pdf", preset="invoice")                    # a path
vision.analyze(file=open("invoice.pdf", "rb"), preset="invoice")        # an open binary file
vision.analyze(file=raw_bytes, preset="invoice")                        # bytes
vision.analyze(file=("scan.png", raw_bytes), preset="invoice")          # bytes + a name
vision.analyze(file_url="https://example.com/invoice.pdf", preset="invoice")
vision.analyze(file_base64=b64, preset="invoice")                       # `data:` prefix optional

JPEG, PNG, WebP, TIFF and PDF, up to 20 MB and 50 pages. The type is detected from magic bytes — the filename is ignored.

Options

Argument Default What it does
preset A catalog name, or "auto" to let the API classify the file first (free).
schema Custom fields, alone or on top of a preset.
schema_name A schema saved in your dashboard. Excludes preset and schema.
pages all PDF page selection, e.g. "1-3,7". You pay for selected pages only.
language_hint auto ISO 639-1 code, e.g. "es".
detail "standard" "high" renders pages at higher resolution. Same cost, slower.
output "json" "text" returns raw OCR text instead of fields.
include_raw_text False Adds full_text, the whole transcription, alongside result.
min_confidence "low" Fields below the level come back None, with confidence preserved.

Custom fields

A schema is a flat dict: each key is a field name, each value describes what to extract. It is compiled before any credit moves, so a bad schema costs nothing.

res = vision.analyze(
    file="invoice.pdf",
    preset="invoice",
    schema={
        # Plain form — the string is the description, type defaults to string.
        "machine_serial": 'Serial number of the machine being invoiced, without the "SN:" prefix',

        # Typed form.
        "total_net": {"type": "number", "description": "Total before tax"},
        "signed_on": {"type": "date", "description": "Date the contract was signed"},
        "is_paid": {"type": "boolean", "description": "Whether the invoice is stamped PAID"},

        # Reserved key: injects fields into every row of the preset's line-item array.
        "line_item": {"lot_number": "The lot number printed on the line, if present"},
    },
)

Field names must match ^[a-z][a-z0-9_]{0,63}$. Types are string (default), number, boolean, date, array and object. A custom name that collides with a preset field is a 422 schema_field_conflict — rename it, or use the preset's own field.

Descriptions are the prompt. "The invoice number exactly as printed, without the #" extracts better than "invoice number". Say what to do when the value is missing or ambiguous if it matters.

Reuse a combination by saving it:

vision.create_schema("our-invoices", preset="invoice", schema={"machine_serial": "…"})
vision.analyze(file="invoice.pdf", schema_name="our-invoices")

Picking a preset

28 presets ship with the API. Fetch the catalog rather than hardcoding field names from memory — presets are versioned, and the catalog is the source of truth:

for p in vision.presets():                    # no API key required
    print(p["name"], p["kind"], p["field_count"])

invoice = vision.preset("invoice")
[f["name"] for f in invoice["fields"]]

Three ways to choose:

# 1. You know what it is.
vision.analyze(file="receipt.jpg", preset="receipt")

# 2. You don't, and you want the data anyway. Classification is free.
res = vision.analyze(file="unknown.pdf", preset="auto")
res["detection"]["preset"]        # what ran
res["detection"]["fallback"]      # True = "shape unknown", not a match
res["detection"]["alternatives"]  # the rest of the ranking, best first

# 3. The *type* is the decision — routing a mixed inbox, or refusing to spend
#    on a 40-page PDF until you know what it is. Far cheaper than extracting.
guess = vision.detect(file="unknown.pdf")
if guess["recommended"] == "invoice" and not guess["fallback"]:
    vision.analyze(file="unknown.pdf", preset="invoice")

detect reads page 1 only, so an image and a 300-page PDF cost the same, and it is metered in batches rather than per call: most calls report credits_used: 0 and an occasional one carries the charge. See pricing for the rate.


Questions instead of fields

Up to 5 questions about one file, priced exactly like an extraction. The questions themselves are free.

res = vision.ask(
    file="photo.jpg",
    questions=["Is there a dog in the image?", "How many people are visible?"],
)

for a in res["answers"]:
    print(a["question"], "→", a["verdict"], a["answer"])

verdict is "yes", "no", "uncertain" (a yes/no question the image does not settle) or "n/a" (not a yes/no question). Branch on it instead of parsing the prose.


Long jobs: async and webhooks

Synchronous requests are killed at 60 seconds with a 504 sync_timeout. Anything that might run longer — a long PDF, detail="high", a batch — belongs on the queue.

# Submit, then poll. wait_for_task handles the loop and the failure case.
task = vision.analyze_and_wait(
    file="contract-80-pages.pdf",
    preset="contract",
    pages="1-50",
    poll_interval=2.0,
    max_wait=900,
    on_poll=lambda t: print(t["status"]),
)
print(task["result"]["parties"]["value"])

# Or submit and walk away — the result comes to you.
ref = vision.analyze_async(
    file="contract.pdf",
    preset="contract",
    webhook_url="https://yourapp.com/hooks/vision",
)

Results stay retrievable for 7 days; after that the task raises ResultExpiredError (metadata survives, the payload does not).

Verifying a delivery

Deliveries are signed. Verify over the raw bytes before parsing — a re-serialized body has different bytes and will not match.

from flask import Flask, request
from visionapi import verify_webhook, WebhookSignatureError

app = Flask(__name__)
SECRET = os.environ["VISION_WEBHOOK_SECRET"]

@app.post("/hooks/vision")
def hook():
    try:
        event = verify_webhook(request.get_data(), request.headers.get("X-Vision-Signature"), SECRET)
    except WebhookSignatureError:
        return "", 400          # never parse an unverified body

    queue.put(event)            # event["event"] is 'task.completed' | 'task.failed'
    return "", 202              # any 2xx is success — ack fast, work afterwards

verify_webhook rejects a bad signature, a malformed header and a timestamp more than 5 minutes old, and accepts a delivery if any v1= part matches — which is what makes a secret rotation seamless. Get the secret from https://app.visionapi.io/dashboard/webhooks. Failed deliveries retry at +1 m, +5 m, +15 m and +40 m, then stop.


Errors

Every failure raises a subclass of VisionAPIError carrying the HTTP status, the stable code, and whatever details the endpoint attached. Branch on the class or on code — never on the message text, which is prose and changes.

from visionapi import (
    InsufficientCreditsError,
    RateLimitError,
    SyncTimeoutError,
    UnsupportedTypeError,
    VisionAPIError,
)

try:
    res = vision.analyze(file="scan.pdf", preset="invoice")
except InsufficientCreditsError as e:
    alert_ops(f"needs {e.required}, has {e.available}")   # never retried — it cannot succeed
except SyncTimeoutError:
    task = vision.analyze_and_wait(file="scan.pdf", preset="invoice")
except UnsupportedTypeError:
    quarantine("not an image or a PDF")
except VisionAPIError as e:
    log.error("vision failed", code=e.code, status=e.status, request_id=e.request_id)
Exception HTTP Codes
InvalidRequestError 400 invalid_request
AuthenticationError 401 invalid_api_key, unauthorized
InsufficientCreditsError 402 insufficient_credits — with .required / .available
PermissionDeniedError 403 forbidden, email_not_verified
NotFoundError 404 task_not_found, schema_not_found
ConflictError 409 conflict
ResultExpiredError 410 result_expired
PayloadTooLargeError 413 file_too_large, page_limit_exceeded
UnsupportedTypeError 415 unsupported_type
UnprocessableError 422 pdf_encrypted, invalid_page_selection, invalid_schema, schema_field_conflict, too_many_questions
RateLimitError 429 rate_limited — with .retry_after
TooManyTasksError 429 too_many_tasks — the per-plan async concurrency cap, with .max_tasks. Subclasses RateLimitError, but is not auto-retried: it clears when one of your tasks finishes
InternalError 500 internal_error — with .request_id
ProviderError 502 provider_error
SyncTimeoutError 504 sync_timeout

UsageError (bad arguments), APIConnectionError / APITimeoutError (the request never got a response) and TaskFailedError / TaskTimeoutError come from the client itself. They are named that way deliberately — shadowing the builtin ConnectionError, TimeoutError and PermissionError in your except clauses would be a nasty surprise.

Retries and idempotency

The client retries 429, 500, 502 and network failures — max_retries=3 by default, with the server's own Retry-After honored on 429 and exponential backoff with jitter elsewhere. Input errors and insufficient_credits are never retried, because they cannot succeed.

Every billable POST is sent with a generated Idempotency-Key, so a retried upload replays the first response instead of paying twice. Supply your own when the caller may retry — a job that re-runs, a queue that redelivers — because a fresh process generates a fresh key:

vision.analyze(file=path, preset="invoice", idempotency_key=f"invoice-{invoice_id}")

Reusing a key with a different payload raises ConflictError, which is the mechanism working: it means the key already stands for something else.


Configuration

vision = VisionAPI(
    api_key=os.environ["VISION_API_KEY"],  # default: $VISION_API_KEY
    base_url="https://api.visionapi.io",   # default; override for a self-hosted deployment
    timeout=120.0,                          # per request, seconds
    max_retries=3,
    auto_idempotency=True,
    headers={"x-trace-id": trace_id},       # sent on every request
)

Every method takes per-call idempotency_key= and timeout=.

TLS certificates. The client uses certifi's CA bundle when it is importable, and the system store otherwise. That is deliberate: a python.org macOS build ships with an empty store until you run Install Certificates.command, and the resulting SSLCertVerificationError looks like an API problem rather than an interpreter one. pip install certifi fixes it; set $VISION_CA_BUNDLE to point at your own root if you are behind a TLS-inspecting proxy.


Account and usage

credits = vision.credits()
credits["balance"], credits["buckets"]
# buckets are spent in order: subscription → rollover → pack → welcome

for record in vision.iter_requests(limit=100):
    print(record["created_at"], record["endpoint"], record["preset"], record["credits_used"])

Usage history is metadata only — never the file, never the extracted values. Uploaded files are never retained: a synchronous request holds yours in memory for the length of the call, and an async request stages it only until the worker finishes with it.


Limits

Same for everyone:

Limit Value
Max file size 20 MB
Max PDF pages per request 50
Sync request timeout 60 s

Per plan:

Limit Free Starter Growth Pro Scale
Requests per minute, per key 10 60 120 300 600
Burst capacity 20 120 240 600 1,200
Concurrent async tasks 1 4 8 16 32
Active API keys per account 1 5 10 20 50
Saved schemas 3 10 25 100 unlimited
Max questions per ask 5 5 5 10 10

The rate-limit bucket is per API key, not per account — splitting a workload across keys splits the limit too. The concurrency cap is per account and does not split that way: over it, an async submission answers 429 too_many_tasks and is charged nothing. Higher limits on paid plans: https://visionapi.io/pricing.


Examples

Runnable scripts in examples/:

File What it shows
analyze.py The smallest useful call, and how to read the result
custom_schema.py Custom fields, line-item injection, saved schemas
detect_then_analyze.py Routing a mixed inbox before spending on extraction
async_batch.py A folder of long PDFs, queued with bounded concurrency
webhook_server.py A verified receiver, with no framework
ask.py Visual Q&A and the verdict field
dataframe.py Line items → pandas DataFrame → CSV
export VISION_API_KEY=sk_live_…
python examples/analyze.py invoice.pdf

Development

pip install -e ".[dev]"
pytest          # offline: the transport is stubbed, no key and no network needed
mypy src
ruff check .

Contributing

Issues and pull requests are welcome at https://github.com/devrobotlabs/visionapi-python. For anything about the API itself — a preset, a limit, an error code — https://support.visionapi.io reaches the team faster.

License

MIT © Vision API

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

visionapi_client-1.0.0.tar.gz (37.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

visionapi_client-1.0.0-py3-none-any.whl (31.5 kB view details)

Uploaded Python 3

File details

Details for the file visionapi_client-1.0.0.tar.gz.

File metadata

  • Download URL: visionapi_client-1.0.0.tar.gz
  • Upload date:
  • Size: 37.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for visionapi_client-1.0.0.tar.gz
Algorithm Hash digest
SHA256 1079b44fbfb495402187be06f51791e7c1c9feeb3873a99d9a245d7c43cf0f4f
MD5 e438dbfb5301db3860159ce9636b305e
BLAKE2b-256 2dcb5055695c04e65dd45a998f3cbd14aa938209842ed59003616a9127f78a76

See more details on using hashes here.

Provenance

The following attestation bundles were made for visionapi_client-1.0.0.tar.gz:

Publisher: publish.yml on devrobotlabs/visionapi-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file visionapi_client-1.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for visionapi_client-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ad2205603227031c6c8f8c36d3d0fde363def3949f373fd390c6f59e14995c6e
MD5 44a231febcf2a3827730578eb044b814
BLAKE2b-256 81c523eada0b52a6d5e6611a3c103e95878c8c42dfa67857b1aacd0bcb0e3182

See more details on using hashes here.

Provenance

The following attestation bundles were made for visionapi_client-1.0.0-py3-none-any.whl:

Publisher: publish.yml on devrobotlabs/visionapi-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page