Skip to main content

polign — Python client for polign_db

A thin Python client for polign with two interchangeable transports:

  • HTTP (polign.Client) — zero dependencies, pure stdlib. Talks JSON to the server's HTTP listener (default :23000).
  • gRPC (polign.GrpcClient) — install with the [grpc] extra. Talks to the gRPC listener (default :23001).

Both expose the same operations with identical semantics: the data plane (put, put_many, get, get_many, list, delete, search) and collection management (create_collection, get_collection, list_collections, delete_collection, verify_collection).

Install

pip install polign             # HTTP client, no dependencies
pip install 'polign[grpc]'     # + gRPC transport (grpcio, protobuf)

Quick start

from polign import Client

client = Client("http://localhost:23000")

# Upsert. Collections are auto-created on first put, inferring their
# dimension from the vector. Values accept lists or numpy arrays.
client.put("docs", "doc-1", embedding, metadata={"title": "Cats", "url": "/cats"})

# Nearest-neighbour search (distance: smaller = closer)
for hit in client.search("docs", values=query_embedding, k=10):
    print(hit.id, hit.distance, hit.metadata)

Swap in gRPC by changing two lines — the rest of the code is identical:

from polign import GrpcClient

client = GrpcClient("localhost:23001")

Operations

client.put("docs", "doc-1", values, metadata={"k": "v"})  # upsert, returns id
client.put_many("docs", [Vector(id="a", values=va), Vector(id="b", values=vb)])
                                               # batch upsert, one request
v = client.get("docs", "doc-1")                # Vector(id, values, metadata)
vs = client.get_many("docs", ["a", "b"])       # batch read, byte-exact values:
                                               # never a compressed reconstruction
                                               # (get may return one on a cold-
                                               # flushed collection); unknown ids
                                               # omitted, request order kept
page = client.list("docs", limit=100, offset=0)  # page.vectors, page.total
page = client.list("docs", filter={"user_id": "u1"})
                                               # filtered listing: same dict
                                               # language as search; offset and
                                               # page.total count matches only,
                                               # so pagination works unchanged
client.delete("docs", "doc-1")                 # True; False if absent (no error)
hits = client.search("docs", values=q, k=10)   # [Hit(id, distance, score, metadata)]

Search options

from polign import Fusion

client.search(
    "docs",
    values=q,                      # vector leg (either values or text required)
    k=10,
    ef=64,                         # HNSW beam width override (0 = server default)
    filter={"lang": "en"},         # metadata predicate (see below)
    text="quick brown fox",        # BM25 leg (needs a segment index server-side)
    fusion=Fusion(method="linear", alpha=0.6),  # hybrid fusion; default RRF
    cold=True, nprobe=8,           # serve from object-store segments
)

text alone runs a pure BM25 search; values + text runs hybrid search fused server-side. hit.score is the BM25/fused relevance (larger = better) and is 0.0 on a pure vector search.

filter takes the same dict language on both transports: bare values are equality (ANDed across keys); per-key operator objects ($eq, $ne, $in, $gt, $gte, $lt, $lte, $exists) and the composers $and/$or/$not express richer predicates:

filter={
    "tenant": "acme",
    "score": {"$gte": 0.5},
    "$or": [{"lang": "en"}, {"lang": {"$exists": False}}],
}

Collection management

Collections are auto-created on first put, so most applications never touch these. On a server started with -byo-store, the collection API additionally binds collections to customer-owned buckets — and it takes an API key:

from polign import CollectionBackend

client = Client("https://db.example.com:23000", api_key="plgn_<key_id>_<secret>")

info = client.create_collection(
    "docs", CollectionBackend(uri="s3://my-bucket/docs", role_arn="arn:aws:iam::…")
)                                     # info.status: "active" or "pending"
info = client.get_collection("docs")  # describe one collection
cols = client.list_collections()      # every registered collection
client.verify_collection("docs")      # re-run bucket verification now
client.delete_collection("docs")      # unregister; bucket data is untouched

A pending collection activates automatically (within ~30s) once you finish your side: write info.claim_token to info.claim_path in the bucket, or attach the trust policy from client.backend_setup(uri) to the role. Without -byo-store these endpoints raise NotEnabledError.

Auth

client = Client(
    "https://db.example.com:23000",
    api_key="plgn_<key_id>_<secret>",     # sent as Authorization: Bearer
)

The data plane takes no credential; the API key guards the collection API on -byo-store servers. With TLS enabled server-side, use an https:// URL (HTTP) or pass credentials=grpc.ssl_channel_credentials() (gRPC).

Errors

All errors subclass polign.PolignError:

Exception HTTP gRPC
InvalidArgumentError 400 INVALID_ARGUMENT
AuthenticationError 401 UNAUTHENTICATED
PermissionDeniedError 403 PERMISSION_DENIED
NotFoundError 404 NOT_FOUND
NotOwnerError 421 FAILED_PRECONDITION
RateLimitError 429 RESOURCE_EXHAUSTED
ServerError 5xx INTERNAL
ConnectionError UNAVAILABLE

NotOwnerError.owner names the owning node in fleet mode — reconnect there and retry.

Notes & caveats

Caveats shared by both transports:

  • Embed documents and queries with the same model — distances are only meaningful within one embedding space.
  • Auto-created collections use the server's default metric (L2) and hybrid IVF index; metric and index tuning are not yet exposed over the wire.
  • Bulk loads should use put_many — one request per batch instead of one per vector. The server validates the whole batch up front (id, non-empty values, uniform dimension, at most 5000 vectors per batch: an invalid batch applies nothing); on a rarer mid-batch failure earlier vectors remain applied, and since puts are idempotent upserts you simply retry the batch. Chunk larger loads into batches of 5000.
  • Metadata is str -> str only; filter scalars (numbers, booleans) compare against the stored string by their literal form ("0.5", "true").

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

polign-0.2.0.tar.gz (27.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

polign-0.2.0-py3-none-any.whl (26.2 kB view details)

Uploaded Python 3

File details

Details for the file polign-0.2.0.tar.gz.

File metadata

  • Download URL: polign-0.2.0.tar.gz
  • Upload date:
  • Size: 27.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for polign-0.2.0.tar.gz
Algorithm Hash digest
SHA256 773c3006a3f7d2b8a88b17ad4da718add63965ba3e7b9fad0a9a5fa5c1ccde36
MD5 a3322c7995eeee97e1373e22a7fb100d
BLAKE2b-256 3a23702c4b43a0fbcd464009c4e264f54451bc1099f396d3c3e11580e86f184b

See more details on using hashes here.

Provenance

The following attestation bundles were made for polign-0.2.0.tar.gz:

Publisher: release.yml on Polign/polign_db

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file polign-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: polign-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 26.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for polign-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e0977916100d2372f0558ebd75c45f8d16feca40597d72f800e00527b5e86cb4
MD5 89ec52e673b7a09fe1ed9b0396017535
BLAKE2b-256 43163fb9b00435572304eeb89b8c488ce7951e68991031af7fd745d8ead8f2f8

See more details on using hashes here.

Provenance

The following attestation bundles were made for polign-0.2.0-py3-none-any.whl:

Publisher: release.yml on Polign/polign_db

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.3.1

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

This release

0.2.0 This release

2 files

0.1.0.post2

2 files

0.1.0.post1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page