Skip to main content

polign — Python client for polign_db

A thin Python client for polign with two interchangeable transports:

  • HTTP (polign.Client) — zero dependencies, pure stdlib. Talks JSON to the server's HTTP listener (default :23000).
  • gRPC (polign.GrpcClient) — install with the [grpc] extra. Wire-compatible with the Go client, against the gRPC listener (default :23001).

Both expose the same seven operations with identical semantics: put, put_many, get, get_many, list, delete, search.

Install

Not yet published to PyPI — install from a checkout of the repo:

pip install ./sdk/python           # HTTP client, no dependencies
pip install './sdk/python[grpc]'   # + gRPC transport (grpcio, protobuf)

(From inside this directory: pip install .; add -e for an editable development install — needs pip ≥ 21.3.)

Quick start

from polign import Client

client = Client("http://localhost:23000")

# Upsert. Collections are auto-created on first put, inferring their
# dimension from the vector. Values accept lists or numpy arrays.
client.put("docs", "doc-1", embedding, metadata={"title": "Cats", "url": "/cats"})

# Nearest-neighbour search (distance: smaller = closer)
for hit in client.search("docs", values=query_embedding, k=10):
    print(hit.id, hit.distance, hit.metadata)

Swap in gRPC by changing two lines — the rest of the code is identical:

from polign import GrpcClient

client = GrpcClient("localhost:23001")

Operations

client.put("docs", "doc-1", values, metadata={"k": "v"})  # upsert, returns id
client.put_many("docs", [Vector(id="a", values=va), Vector(id="b", values=vb)])
                                               # batch upsert, one request
v = client.get("docs", "doc-1")                # Vector(id, values, metadata)
vs = client.get_many("docs", ["a", "b"])       # batch read, byte-exact values:
                                               # never a compressed reconstruction
                                               # (get may return one on a cold-
                                               # flushed collection); unknown ids
                                               # omitted, request order kept
page = client.list("docs", limit=100, offset=0)  # page.vectors, page.total
client.delete("docs", "doc-1")                 # True; False if absent (no error)
hits = client.search("docs", values=q, k=10)   # [Hit(id, distance, score, metadata)]

Search options

from polign import Fusion

client.search(
    "docs",
    values=q,                      # vector leg (either values or text required)
    k=10,
    ef=64,                         # HNSW beam width override (0 = server default)
    filter={"lang": "en"},         # metadata predicate (see below)
    text="quick brown fox",        # BM25 leg (needs a segment index server-side)
    fusion=Fusion(method="linear", alpha=0.6),  # hybrid fusion; default RRF
    cold=True, nprobe=8,           # serve from object-store segments
)

text alone runs a pure BM25 search; values + text runs hybrid search fused server-side. hit.score is the BM25/fused relevance (larger = better) and is 0.0 on a pure vector search.

filter takes the same dict language on both transports (see docs/FILTERING.md in the main repo): bare values are equality (ANDed across keys); per-key operator objects ($eq, $ne, $in, $gt, $gte, $lt, $lte, $exists) and the composers $and/$or/$not express richer predicates:

filter={
    "tenant": "acme",
    "score": {"$gte": 0.5},
    "$or": [{"lang": "en"}, {"lang": {"$exists": False}}],
}

Auth and tenants

client = Client(
    "https://db.example.com:23000",
    api_key="plgn_<key_id>_<secret>",     # sent as Authorization: Bearer
    tenant="acme/search/prod",            # org/project/namespace
)

Servers started without -auth-stores need no credentials. With TLS enabled server-side, use an https:// URL (HTTP) or pass credentials=grpc.ssl_channel_credentials() (gRPC).

Errors

All errors subclass polign.PolignError:

Exception HTTP gRPC
InvalidArgumentError 400 INVALID_ARGUMENT
AuthenticationError 401 UNAUTHENTICATED
PermissionDeniedError 403 PERMISSION_DENIED
NotFoundError 404 NOT_FOUND
NotOwnerError 421 FAILED_PRECONDITION
RateLimitError 429 RESOURCE_EXHAUSTED
ServerError 5xx INTERNAL
ConnectionError UNAVAILABLE

NotOwnerError.owner names the owning node in fleet mode — reconnect there and retry.

Notes & caveats

Mirrors the Go client's caveats (see docs/CLIENT.md in the main repo):

  • Embed documents and queries with the same model — distances are only meaningful within one embedding space.
  • Auto-created collections use the server's default metric (L2) and hybrid IVF index; metric and index tuning are not yet exposed over the wire.
  • Bulk loads should use put_many — one request per batch instead of one per vector. The server validates the whole batch up front (id, non-empty values, uniform dimension, at most 5000 vectors per batch: an invalid batch applies nothing); on a rarer mid-batch failure earlier vectors remain applied, and since puts are idempotent upserts you simply retry the batch. Chunk larger loads into batches of 5000.
  • Metadata is str -> str only; filter scalars (numbers, booleans) compare against the stored string by their literal form ("0.5", "true").

Development

cd sdk/python
pip install -e .[grpc] pytest
pytest tests/test_unit.py          # stub-server unit tests
pytest tests/test_integration.py   # builds & boots the real server (needs Go)

Regenerate the vendored gRPC stubs after changing proto/vectordb.proto (from the repo root):

python -m grpc_tools.protoc -I proto \
    --python_out=sdk/python/polign/_pb \
    --grpc_python_out=sdk/python/polign/_pb \
    --pyi_out=sdk/python/polign/_pb \
    proto/vectordb.proto
sed -i '' 's/^import vectordb_pb2 as/from . import vectordb_pb2 as/' \
    sdk/python/polign/_pb/vectordb_pb2_grpc.py

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

polign-0.1.0.tar.gz (26.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

polign-0.1.0-py3-none-any.whl (25.8 kB view details)

Uploaded Python 3

File details

Details for the file polign-0.1.0.tar.gz.

File metadata

  • Download URL: polign-0.1.0.tar.gz
  • Upload date:
  • Size: 26.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for polign-0.1.0.tar.gz
Algorithm Hash digest
SHA256 3fb50dfdf0b083480b24292eb1d17aa39df7e58b3f67f0e17da2fbfd060f80e0
MD5 2dc39d3085d42636e719366f168646c6
BLAKE2b-256 3009a16e09c552368e68d15a599e3a8376240c92f199dcb5846983152f7c4104

See more details on using hashes here.

File details

Details for the file polign-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: polign-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 25.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for polign-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 972e1ce4afa789ae89289f29cd30346890b8346d3c972182543c8cd49e586c02
MD5 fcd5c799e3e794f4538717d1ca19cf0e
BLAKE2b-256 b6cf8d4570ad88da25bf9dc6154dd224669516345f1b6c532d88ed6dae8d484b

See more details on using hashes here.

Release history Release notifications | RSS feed

0.3.1

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0.post2

2 files

0.1.0.post1

2 files

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page