Skip to main content

navia-sdk

Official Python SDK for the Navia platform.

The SDK lets external services talk to an Navia backend over HTTP. A single class — Navia — wraps every endpoint the backend publishes under /sdk/*:

Capability Token permission What it does
External connectors connector Register a connector, run sync cycles, push files, query checksums.
Artifacts mcp Publish generated files for delivery in chat.
OCR ocr Extract Markdown/JSON/PDF from a document.
Data extraction data-extraction Extract structured data from a document against a saved template or inline schema.
Chat chat OpenAI-style chat completions (not implemented yet).

A token's permissions decide which methods are callable; the backend rejects calls made with a token that lacks the required permission. The chat methods raise NotImplementedError until the backend implementation lands.

Installation

pip install navia-sdk

Imports use the top-level navia package:

from navia import Navia

Authentication

All calls authenticate with an API token sent through the X-Api-Key HTTP header. Tokens are created from the Navia admin panel and carry one or more permissions (connector, mcp, ocr, chat).

The client reads two environment variables by default (or accepts them as constructor arguments):

Variable Purpose
NAVIA_API_URL Base URL of the Navia backend.
NAVIA_API_TOKEN API token (omni_...).
from navia import Navia

# Reads NAVIA_API_URL / NAVIA_API_TOKEN from the environment.
with Navia() as client:
    print(client.connector_id)        # connector bound to this token, if any
    print(client.last_sync_started_at)  # checkpoint of the last sync run

# …or pass them explicitly:
client = Navia(api_url="https://navia.example.com", api_token="omni_…")

When the token is bound to a connector, the client resolves the connector ID on construction. Pass resolve_connector=False to skip that lookup when you only use the MCP or OCR features.

External connectors

A connector pushes documents from an external source into Navia. The run_sync helper drives a full cycle — optional auto-register → notify start → your sync function → notify end (with stale-item cleanup):

import os
from pathlib import Path

from navia import Navia


def sync(client: Navia) -> list[str]:
    source_dir = Path(os.environ["NAVIA_SOURCE_DIR"])
    existing = set(client.get_existing_checksums())
    active: list[str] = []

    for path in sorted(source_dir.rglob("*")):
        if not path.is_file():
            continue
        checksum = Navia.compute_checksum(path)
        active.append(checksum)
        if checksum not in existing:
            client.push_file(path, source_id=str(path.relative_to(source_dir)))

    # Returning the active checksums lets the backend delete stale items.
    return active


if __name__ == "__main__":
    with Navia() as client:
        client.run_sync(sync, name="local-files", description="Local files")

item_exists(source_id=...) / item_exists(checksum=...) query the backend for a single item without fetching the whole checksum list. See examples/external_connector/ for a runnable connector and Dockerfile.

Artifacts

Publish a file produced by an MCP tool as an Artifact, and fetch it back by id:

from navia import Navia

with Navia(resolve_connector=False) as client:
    info = client.artifact_upload("./report.pdf", display_name="Q4 report")
    client.artifact_download(info["artifactId"], "./downloaded.pdf")

Uploading alone only stages the file. To deliver it to the chat user, the MCP tool result must declare the id under the reserved navia_artifacts key of its structured content:

{
  "navia_artifacts": [
    {"artifact_id": "<artifactId>", "filename": "report.pdf",
     "content_type": "application/pdf", "size": 12345}
  ]
}

Only artifact_id is required — the other fields are hints. The chat backend then shows the file as a downloadable card on the assistant's message and in the user's Artifacts list. Identical re-uploads by the same token are deduplicated server-side (deduped: true in the response).

mcp_upload_attachment / mcp_download_attachment are deprecated aliases of the old /sdk/mcp/* routes and will be removed in 0.5.0.

OCR

Run OCR on a single document. By default the structured result is returned as a dict; pass output_format together with dest to download the rendered file instead:

from navia import Navia

with Navia(resolve_connector=False) as client:
    result = client.ocr_extract("./document.pdf", mode="STRUCTURED")
    print(result["markdown"])

    # Render and download a file:
    client.ocr_extract(
        "./document.pdf", output_format="MARKDOWN", dest="./document.md"
    )

mode accepts "PLAIN", "STRUCTURED" (default) or "VLM"; output_format accepts "MARKDOWN", "JSON" or "PDF". Result keys are snake_case (markdown, regions, page_count, …).

Data extraction

Extract structured data from a document against a saved template or an inline definition (provide exactly one). Non-PDF inputs are converted to PDF server-side. The structured result is returned as a dict; pass output_format ("csv"/"xlsx") together with dest to download a table:

from navia import Navia

with Navia(resolve_connector=False) as client:
    # Against a saved template:
    result = client.extraction_extract("./invoice.pdf", template_id="<id>")
    for field in result["fields"]:
        print(field["field_path"], field["value"], field["confidence"])

    # Against an inline definition + a hint:
    client.extraction_extract(
        "./invoice.pdf",
        definition={"fields": [{"key": "total", "label": "Total", "type": "currency"}]},
        hints="the grand total is bottom-right",
    )

    # Download a CSV of the extracted fields:
    client.extraction_extract(
        "./invoice.pdf", template_id="<id>", output_format="csv", dest="./out.csv"
    )

Result keys are snake_case: data (the structured record), fields (each with field_path, value, value_type, confidence, source_spans) and usage.

Chat (not implemented)

from navia import Navia

with Navia(resolve_connector=False) as client:
    # Raises NotImplementedError today.
    client.chat_completions({"messages": [{"role": "user", "content": "Hi"}]})

Development

The project uses uv and ruff (pinned in pyproject.toml).

uv sync
uv run ruff check src tests
uv run ty check
uv run pytest

Versioning & release

The package version lives in [project].version of pyproject.toml (a single source of truth; navia.__version__ is read from the installed package metadata). Releases are driven by CI:

  • pushes to develop publish a dev build to TestPyPI (the version is suffixed with .devN so each build is unique);
  • pushes to main publish the exact pyproject.toml version to PyPI as navia-sdk.

Bump version in pyproject.toml before promoting a release to main.

Release files for navia-sdk 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for navia-sdk 2.0.0
File Size Uploaded
navia_sdk-2.0.0.tar.gz 9.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for navia-sdk 2.0.0
File Interpreter ABI Platform
navia_sdk-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 20.9 kB

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page