navia-sdk
Official Python SDK for the Navia platform.
The SDK lets external services talk to an Navia backend over HTTP. A
single class — Navia — wraps every endpoint the backend publishes
under /sdk/*:
| Capability | Token permission | What it does |
|---|---|---|
| External connectors | connector |
Register a connector, run sync cycles, push files, query checksums. |
| Artifacts | mcp |
Publish generated files for delivery in chat. |
| OCR | ocr |
Extract Markdown/JSON/PDF from a document. |
| Data extraction | data-extraction |
Extract structured data from a document against a saved template or inline schema. |
| Chat | chat |
OpenAI-style chat completions (not implemented yet). |
A token's permissions decide which methods are callable; the backend rejects
calls made with a token that lacks the required permission. The chat methods
raise NotImplementedError until the backend implementation lands.
Installation
pip install navia-sdk
Imports use the top-level navia package:
from navia import Navia
Authentication
All calls authenticate with an API token sent through the X-Api-Key HTTP
header. Tokens are created from the Navia admin panel and carry one or
more permissions (connector, mcp, ocr, chat).
The client reads two environment variables by default (or accepts them as constructor arguments):
| Variable | Purpose |
|---|---|
NAVIA_API_URL |
Base URL of the Navia backend. |
NAVIA_API_TOKEN |
API token (omni_...). |
from navia import Navia
# Reads NAVIA_API_URL / NAVIA_API_TOKEN from the environment.
with Navia() as client:
print(client.connector_id) # connector bound to this token, if any
print(client.last_sync_started_at) # checkpoint of the last sync run
# …or pass them explicitly:
client = Navia(api_url="https://navia.example.com", api_token="omni_…")
When the token is bound to a connector, the client resolves the connector ID
on construction. Pass resolve_connector=False to skip that lookup when you
only use the MCP or OCR features.
External connectors
A connector pushes documents from an external source into Navia. The
run_sync helper drives a full cycle — optional auto-register → notify start
→ your sync function → notify end (with stale-item cleanup):
import os
from pathlib import Path
from navia import Navia
def sync(client: Navia) -> list[str]:
source_dir = Path(os.environ["NAVIA_SOURCE_DIR"])
existing = set(client.get_existing_checksums())
active: list[str] = []
for path in sorted(source_dir.rglob("*")):
if not path.is_file():
continue
checksum = Navia.compute_checksum(path)
active.append(checksum)
if checksum not in existing:
client.push_file(path, source_id=str(path.relative_to(source_dir)))
# Returning the active checksums lets the backend delete stale items.
return active
if __name__ == "__main__":
with Navia() as client:
client.run_sync(sync, name="local-files", description="Local files")
item_exists(source_id=...) / item_exists(checksum=...) query the backend
for a single item without fetching the whole checksum list. See
examples/external_connector/ for a runnable
connector and Dockerfile.
Artifacts
Publish a file produced by an MCP tool as an Artifact, and fetch it back by id:
from navia import Navia
with Navia(resolve_connector=False) as client:
info = client.artifact_upload("./report.pdf", display_name="Q4 report")
client.artifact_download(info["artifactId"], "./downloaded.pdf")
Uploading alone only stages the file. To deliver it to the chat user, the
MCP tool result must declare the id under the reserved
navia_artifacts key of its structured content:
{
"navia_artifacts": [
{"artifact_id": "<artifactId>", "filename": "report.pdf",
"content_type": "application/pdf", "size": 12345}
]
}
Only artifact_id is required — the other fields are hints. The chat backend
then shows the file as a downloadable card on the assistant's message and in
the user's Artifacts list. Identical re-uploads by the same token are
deduplicated server-side (deduped: true in the response).
mcp_upload_attachment/mcp_download_attachmentare deprecated aliases of the old/sdk/mcp/*routes and will be removed in 0.5.0.
OCR
Run OCR on a single document. By default the structured result is returned as
a dict; pass output_format together with dest to download the rendered
file instead:
from navia import Navia
with Navia(resolve_connector=False) as client:
result = client.ocr_extract("./document.pdf", mode="STRUCTURED")
print(result["markdown"])
# Render and download a file:
client.ocr_extract(
"./document.pdf", output_format="MARKDOWN", dest="./document.md"
)
mode accepts "PLAIN", "STRUCTURED" (default) or "VLM";
output_format accepts "MARKDOWN", "JSON" or "PDF". Result keys are
snake_case (markdown, regions, page_count, …).
Data extraction
Extract structured data from a document against a saved template or an
inline definition (provide exactly one). Non-PDF inputs are converted to
PDF server-side. The structured result is returned as a dict; pass
output_format ("csv"/"xlsx") together with dest to download a table:
from navia import Navia
with Navia(resolve_connector=False) as client:
# Against a saved template:
result = client.extraction_extract("./invoice.pdf", template_id="<id>")
for field in result["fields"]:
print(field["field_path"], field["value"], field["confidence"])
# Against an inline definition + a hint:
client.extraction_extract(
"./invoice.pdf",
definition={"fields": [{"key": "total", "label": "Total", "type": "currency"}]},
hints="the grand total is bottom-right",
)
# Download a CSV of the extracted fields:
client.extraction_extract(
"./invoice.pdf", template_id="<id>", output_format="csv", dest="./out.csv"
)
Result keys are snake_case: data (the structured record), fields (each
with field_path, value, value_type, confidence, source_spans) and
usage.
Chat (not implemented)
from navia import Navia
with Navia(resolve_connector=False) as client:
# Raises NotImplementedError today.
client.chat_completions({"messages": [{"role": "user", "content": "Hi"}]})
Development
The project uses uv and
ruff (pinned in pyproject.toml).
uv sync
uv run ruff check src tests
uv run ty check
uv run pytest
Versioning & release
The package version lives in [project].version of pyproject.toml (a single
source of truth; navia.__version__ is read from the installed package
metadata). Releases are driven by CI:
- pushes to
developpublish a dev build to TestPyPI (the version is suffixed with.devNso each build is unique); - pushes to
mainpublish the exactpyproject.tomlversion to PyPI asnavia-sdk.
Bump version in pyproject.toml before promoting a release to main.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters