Skip to main content

polign — Python client for polign_db

A thin Python client for polign with two interchangeable transports:

  • HTTP (polign.Client) — zero dependencies, pure stdlib. Talks JSON to the server's HTTP listener (default :23000).
  • gRPC (polign.GrpcClient) — install with the [grpc] extra. Talks to the gRPC listener (default :23001).

Both expose the same operations with identical semantics: the data plane (put, put_many, get, get_many, list, delete, search) and collection management (create_collection, get_collection, list_collections, delete_collection, verify_collection).

Install

pip install polign             # HTTP client, no dependencies
pip install 'polign[grpc]'     # + gRPC transport (grpcio, protobuf)

Quick start

from polign import Client

client = Client("http://localhost:23000")

# Upsert. Collections are auto-created on first put, inferring their
# dimension from the vector. Values accept lists or numpy arrays.
client.put("docs", "doc-1", embedding, metadata={"title": "Cats", "url": "/cats"})

# Nearest-neighbour search (distance: smaller = closer)
for hit in client.search("docs", values=query_embedding, k=10):
    print(hit.id, hit.distance, hit.metadata)

Swap in gRPC by changing two lines — the rest of the code is identical:

from polign import GrpcClient

client = GrpcClient("localhost:23001")

Operations

client.put("docs", "doc-1", values, metadata={"k": "v"})  # upsert, returns id
client.put_many("docs", [Vector(id="a", values=va), Vector(id="b", values=vb)])
                                               # batch upsert, one request
v = client.get("docs", "doc-1")                # Vector(id, values, metadata)
vs = client.get_many("docs", ["a", "b"])       # batch read, byte-exact values:
                                               # never a compressed reconstruction
                                               # (get may return one on a cold-
                                               # flushed collection); unknown ids
                                               # omitted, request order kept
page = client.list("docs", limit=100, offset=0)  # page.vectors, page.total
page = client.list("docs", filter={"user_id": "u1"})
                                               # filtered listing: same dict
                                               # language as search; offset and
                                               # page.total count matches only,
                                               # so pagination works unchanged
client.delete("docs", "doc-1")                 # True; False if absent (no error)
hits = client.search("docs", values=q, k=10)   # [Hit(id, distance, score, metadata)]

Typed metadata

Metadata values may be strings, numbers, or booleans; numbers and booleans are stored typed, and filters compare them by type ({"score": {"$gt": 0.5}} matches numerically). By default reads return every value as a string, so existing code keeps working; pass typed_metadata=True to get, get_many, list, or search to get stored types back:

client.put("docs", "m1", values, metadata={"topic": "SIP", "score": 0.85, "published": True})
v = client.get("docs", "m1")                        # {"score": "0.85", ...}  (strings)
v = client.get("docs", "m1", typed_metadata=True)   # {"score": 0.85, "published": True, ...}
hits = client.search("docs", values=q, k=10, filter={"score": {"$gt": 0.5}})

Typed values need a server with typed-metadata support (v0.3.0+); all-string metadata works against any server version.

Search options

from polign import Fusion

client.search(
    "docs",
    values=q,                      # vector leg (either values or text required)
    k=10,
    ef=64,                         # HNSW beam width override (0 = server default)
    filter={"lang": "en"},         # metadata predicate (see below)
    text="quick brown fox",        # BM25 leg (needs a segment index server-side)
    fusion=Fusion(method="linear", alpha=0.6),  # hybrid fusion; default RRF
    cold=True, nprobe=8,           # serve from object-store segments
)

text alone runs a pure BM25 search; values + text runs hybrid search fused server-side. hit.score is the BM25/fused relevance (larger = better) and is 0.0 on a pure vector search.

filter takes the same dict language on both transports: bare values are equality (ANDed across keys); per-key operator objects ($eq, $ne, $in, $gt, $gte, $lt, $lte, $exists) and the composers $and/$or/$not express richer predicates:

filter={
    "tenant": "acme",
    "score": {"$gte": 0.5},
    "$or": [{"lang": "en"}, {"lang": {"$exists": False}}],
}

Collection management

Collections are auto-created on first put, so most applications never touch these. On a server started with -byo-store, the collection API additionally binds collections to customer-owned buckets — and it takes an API key:

from polign import CollectionBackend

client = Client("https://db.example.com:23000", api_key="plgn_<key_id>_<secret>")

info = client.create_collection(
    "docs", CollectionBackend(uri="s3://my-bucket/docs", role_arn="arn:aws:iam::…")
)                                     # info.status: "active" or "pending"
info = client.get_collection("docs")  # describe one collection
cols = client.list_collections()      # every registered collection
client.verify_collection("docs")      # re-run bucket verification now
client.delete_collection("docs")      # unregister; bucket data is untouched

A pending collection activates automatically (within ~30s) once you finish your side: write info.claim_token to info.claim_path in the bucket, or attach the trust policy from client.backend_setup(uri) to the role. Without -byo-store these endpoints raise NotEnabledError.

Auth

client = Client(
    "https://db.example.com:23000",
    api_key="plgn_<key_id>_<secret>",     # sent as Authorization: Bearer
)

The data plane takes no credential; the API key guards the collection API on -byo-store servers. With TLS enabled server-side, use an https:// URL (HTTP) or pass credentials=grpc.ssl_channel_credentials() (gRPC).

Errors

All errors subclass polign.PolignError:

Exception HTTP gRPC
InvalidArgumentError 400 INVALID_ARGUMENT
AuthenticationError 401 UNAUTHENTICATED
PermissionDeniedError 403 PERMISSION_DENIED
NotFoundError 404 NOT_FOUND
NotOwnerError 421 FAILED_PRECONDITION
RateLimitError 429 RESOURCE_EXHAUSTED
ServerError 5xx INTERNAL
ConnectionError UNAVAILABLE

NotOwnerError.owner names the owning node in fleet mode — reconnect there and retry.

Notes & caveats

Caveats shared by both transports:

  • Embed documents and queries with the same model — distances are only meaningful within one embedding space.
  • Auto-created collections use the server's default metric (L2) and hybrid IVF index; metric and index tuning are not yet exposed over the wire.
  • Bulk loads should use put_many — one request per batch instead of one per vector. The server validates the whole batch up front (id, non-empty values, uniform dimension, at most 5000 vectors per batch: an invalid batch applies nothing); on a rarer mid-batch failure earlier vectors remain applied, and since puts are idempotent upserts you simply retry the batch. Chunk larger loads into batches of 5000.
  • Metadata is str -> str only; filter scalars (numbers, booleans) compare against the stored string by their literal form ("0.5", "true").

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

polign-0.3.1.tar.gz (29.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

polign-0.3.1-py3-none-any.whl (28.0 kB view details)

Uploaded Python 3

File details

Details for the file polign-0.3.1.tar.gz.

File metadata

  • Download URL: polign-0.3.1.tar.gz
  • Upload date:
  • Size: 29.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for polign-0.3.1.tar.gz
Algorithm Hash digest
SHA256 b9247f877e7beb6f0aac364afea22bffbf929f1ac88b034621ee0b04dc3c497e
MD5 a7cf35e58c06aa2e40a0baad44423c22
BLAKE2b-256 f7cf460135cfd8db56bcc65b35cc43328a3695dfd120b3700688bca3b9f2b765

See more details on using hashes here.

Provenance

The following attestation bundles were made for polign-0.3.1.tar.gz:

Publisher: release.yml on Polign/polign_db

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file polign-0.3.1-py3-none-any.whl.

File metadata

  • Download URL: polign-0.3.1-py3-none-any.whl
  • Upload date:
  • Size: 28.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for polign-0.3.1-py3-none-any.whl
Algorithm Hash digest
SHA256 900347e1e2ca8442a6a801334957e25617f68ff8ab297a05c4dad7d919ad0d9f
MD5 51056ffad8deedf5e8b6cf019ab1b800
BLAKE2b-256 15981392ca6d4a3a5656f5a3893cd7da6205eaf69ee8709167e1d6e8a4b1dd68

See more details on using hashes here.

Provenance

The following attestation bundles were made for polign-0.3.1-py3-none-any.whl:

Publisher: release.yml on Polign/polign_db

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0.post2

2 files

0.1.0.post1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page