Skip to main content

PaveDB Python SDK

Python SDK package for the PaveDB /v1 API.

Use pavedb-sdk when your code should talk to PaveDB from Python.

There are three runtime paths:

  • Connect to a PaveDB server over HTTP.
  • Install pavedb alongside the SDK and use the same Client / Collection handle API with a local embedded engine.
  • Use ephemeral local mode for temporary in-process stores during tests, notebooks, and short-lived experiments.

To run your own server instance, use the PaveDB core repository: GitLab / GitHub. The core repository remains the source of truth for the OpenAPI contract.

SDK source lives on GitLab and GitHub.

Tenant administration (PaveDB 1.0)

An admin client can create_tenant, get_tenant, update_tenant, and delete_tenant, plus create_tenant_key, list_tenant_keys, and revoke_tenant_key. Generated key plaintext is returned only at creation. Tenant quota overrides accept 0 (none), -1 (unlimited), or None (inherit).

Install

pip install pavedb-sdk

SDK 0.1.x targets the PaveDB /v1 API. SDK package versions are independent from PaveDB core release versions; use pavesdk.__version__ for the SDK release and pavesdk.PAVEDB_API_PREFIX for the wire API.

For local embedded/persisted mode:

pip install pavedb-sdk pavedb

Local Package Build

Build the local PyPI package artifacts with GNU Make:

gmake package

That creates the source distribution (.tar.gz) and wheel in dist/, checks them with Twine, and copies them to artifacts/.

Upload targets are explicit and do not infer release channels from the version:

gmake pypitest-push
gmake pypi-push

Runnable HTTP examples are installed with the package:

python -m pavesdk.examples.http_search
python -m pavesdk.examples.observability

The source distribution also includes the book companion programs under examples/. They are intentionally not installed in the wheel; start with examples/1-intuition/README.md after installing pavedb alongside the SDK.

The SDK checkout also includes demo/20k_leagues.txt, which the examples use via a hardcoded relative path.

Generated API Reference

The source distribution includes its generated API reference, examples index, and generator. From a checkout or unpacked SDK sdist, regenerate them with:

python docs/generate_reference.py

gmake docs-check verifies that the checked-in Markdown is current, every indexed example imports, and each one keeps its python -m pavesdk.examples... command.

HTTP Client

from pavesdk.client import connect

db = connect(
    "http://localhost:8086",
    api_key="super-sekret",
    tenant="demo",
)
books = db.collection("books")

hits = books.search("captain nemo", k=3)
hits
[
    {
        "id": "note-1:0000",
        "score": 0.86,
        "text": "Captain Nemo commands the Nautilus.",
        "meta": {"docid": "note-1", "kind": "note"},
    },
    {
        "id": "note-2:0000",
        "score": 0.73,
        "text": "The Nautilus dives beneath the ice.",
        "meta": {"docid": "note-2", "kind": "note"},
    },
]
for hit in hits:
    print(hit["score"], hit["meta"]["docid"], hit["text"][:80])
0.86 note-1 Captain Nemo commands the Nautilus.
0.73 note-2 The Nautilus dives beneath the ice.

connect("http://...") and connect("https://...") create an HttpClient. Bare paths are local targets and require pavedb to be installed.

Collections

The API is handle-based: pick a collection once, then call methods on it.

from pavesdk.client import connect

db = connect("http://localhost:8086", api_key="super-sekret")
books = db.create_collection("books", tenant="demo")

books.add(
    "Captain Nemo commands the Nautilus.",
    docid="note-1",
    metadata={"kind": "note"},
)
books.add_many([
    ("The Nautilus dives beneath the ice.", "note-2", None),
    {
        "text": "Nemo studies ocean currents.",
        "docid": "note-3",
        "metadata": {"kind": "note"},
    },
])

matches = books.search(
    "submarine captain",
    k=5,
    filters={"kind": "note"},
)
matches
[
    {
        "id": "note-1:0000",
        "score": 0.81,
        "text": "Captain Nemo commands the Nautilus.",
        "meta": {"docid": "note-1", "kind": "note"},
    },
    {
        "id": "note-3:0000",
        "score": 0.69,
        "text": "Nemo studies ocean currents.",
        "meta": {"docid": "note-3", "kind": "note"},
    },
]

Observability

Searches are logged by PaveDB. Use query inspection to see what ran, replay it against current data, and inspect the source chunks behind a document.

from pavesdk.client import connect

db = connect("http://localhost:8086", api_key="super-sekret")
books = db.collection("books", tenant="demo")

books.search("captain nemo", k=3)

latest = books.queries(limit=1)[0]
latest
{
    "query_id": "0d4f5a1b-9e4b-41c7-8b3f-8f6b5de3e74a",
    "tenant": "demo",
    "collection": "books",
    "query_text": "captain nemo",
    "k": 3,
    "filters": None,
    "result_count": 2,
    "latency_ms": 12.4,
    "created_at": "2026-06-20T18:42:16.153201Z",
}
query = books.get_query(latest["query_id"])
query
{
    "query_id": "0d4f5a1b-9e4b-41c7-8b3f-8f6b5de3e74a",
    "tenant": "demo",
    "collection": "books",
    "query_text": "captain nemo",
    "k": 3,
    "filters": None,
    "result_ids": ["note-1:0000", "note-2:0000"],
    "result_count": 2,
    "latency_ms": 12.4,
}
replayed = books.replay(query["query_id"])
replayed
[
    {
        "id": "note-1:0000",
        "score": 0.86,
        "text": "Captain Nemo commands the Nautilus.",
        "meta": {"docid": "note-1", "kind": "note"},
    },
    {
        "id": "note-2:0000",
        "score": 0.73,
        "text": "The Nautilus dives beneath the ice.",
        "meta": {"docid": "note-2", "kind": "note"},
    },
]
docid = replayed[0]["meta"]["docid"]
chunks = books.list_chunks(docid)
chunks
[
    {
        "rid": "note-1:0000",
        "docid": "note-1",
        "chunk": 0,
        "text": "Captain Nemo commands the Nautilus.",
        "metadata": {"kind": "note"},
    }
]
chunk = books.get_chunk(chunks[0]["rid"])
chunk
{
    "rid": "note-1:0000",
    "docid": "note-1",
    "chunk": 0,
    "text": "Captain Nemo commands the Nautilus.",
    "metadata": {"kind": "note"},
}
content = books.get_chunk_content(chunk["rid"])
content
{
    "content": b"Captain Nemo commands the Nautilus.",
    "content_type": "text/plain; charset=utf-8",
}

Local Mode

With pavedb installed, the same API can use a local persisted store:

from pavesdk.client import connect

with connect("./data", tenant="demo") as db:
    books = db.create_collection("books")
    books.add("Captain Nemo commands the Nautilus.", docid="note-1")
    print(books.search("captain", k=3))

If pavedb is not installed, local targets raise LocalClientUnavailable.

Archives

from pathlib import Path
from pavesdk.client import connect

with connect("http://localhost:8086", api_key="super-sekret") as db:
    archive_bytes = db.dump_archive()
    Path("pavedb-data.zip").write_bytes(archive_bytes)

    saved_path = db.dump_archive("pavedb-data.zip")
    db.restore_archive(Path(saved_path).read_bytes())

One collection moves on its own. restore_archive on a new name creates the collection; replace=True rolls an existing one back to the snapshot:

with connect("http://localhost:8086", api_key="tenant-key", tenant="demo") as db:
    books = db.collection("books")
    snapshot = books.dump_archive()
    db.collection("books-copy").restore_archive(snapshot)
    books.restore_archive(snapshot, replace=True)

    job = books.reindex(embed_model="sentence-transformers/all-MiniLM-L6-v2")
    print(books.reindex_job(job["job_id"])["status"])

Metadata

Release files for pavedb-sdk 0.1.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pavedb-sdk 0.1.6
File Size Uploaded
pavedb_sdk-0.1.6.tar.gz 273.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pavedb-sdk 0.1.6
File Interpreter ABI Platform
pavedb_sdk-0.1.6-py3-none-any.whl Python 3 none any Details

Total release size: 294.8 kB

Release files / pavedb_sdk-0.1.6.tar.gz

Download URL pavedb_sdk-0.1.6.tar.gz
Size 273.9 kB
Tags Source
SHA-256 checksum
How to use checksums
9314839764160259bd9657bec334cdcffdcfa7f3e77b6739a84921a36be09e9f
BLAKE2b-256 checksum
How to use checksums
5202e255250607d860b263e752373d08d240c22a46dfc1a7136d1cda47d46a86
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / pavedb_sdk-0.1.6-py3-none-any.whl

Download URL pavedb_sdk-0.1.6-py3-none-any.whl
Size 20.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
77c2b72fe1c16a0db4658e485bb99cf411202a9a97fa30ffeb33e5b05d441b7e
BLAKE2b-256 checksum
How to use checksums
eaa38bce8be9b9c4b8c82e99a2d11c553212fbee7d075a2cc1e68aff10e52b7f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

This release

0.1.6 This release

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page