Skip to main content

nlweb-goodmem

GoodMem as an NLWeb retrieval provider.

pip install nlweb-goodmem

Configure

NLWeb imports a provider by path, so this package never has to be added to NLWeb's own source tree:

retrieval:
  default:
    import_path: nlweb_goodmem
    class_name: GoodMemRetrievalProvider
    base_url: https://localhost:8080
    api_key: gm_…
    space_name: nlweb

NLWeb's ask handler asks for the retrieval provider named default, so that is the name to give it. Options sit beside import_path and class_name: NLWeb passes every other key of the entry to the constructor as a keyword argument. A key ending in _env is read from the environment instead, so api_key_env: GOODMEM_API_KEY keeps the key out of the file.

It implements nlweb_core.retriever.RetrievalProvider, so search() returns real RetrievedItem objects and close() releases the client.

space_id, embedder_id and reranker_id must be GoodMem ids, which are UUIDs; any other value raises ValueError when the provider is built, before a request is made. (An empty embedder_id or reranker_id means "not set"; an empty space_id is refused.) The GoodMem SDK places ids in the URL path unencoded, so a value like ../spaces/<id> would otherwise send the request to a different endpoint.

Sites

NLWeb filters every call by site; GoodMem has no such concept. A site is stored as memory metadata and filtered server-side with GoodMem's filter grammar, the same way NLWeb's own Qdrant provider uses a site payload field. Values are escaped, never interpolated — a site called x' OR '1'='1 matches nothing rather than everything.

await provider.search("noodle soup", "recipes.example.com")   # one site
await provider.search("noodle soup", ["a.com", "b.com"])      # OR across sites
await provider.search("noodle soup", "all")                   # every site
await provider.search_all_sites("noodle soup")                # same thing
await provider.get_sites()                                    # ['a.com', 'b.com']

Ingesting Schema.org documents

RetrievalProvider only reads, so ingestion is a helper rather than part of the interface:

from nlweb_goodmem import GoodMemRetrievalProvider, upload_documents

provider = GoodMemRetrievalProvider(space_name="nlweb", base_url=…, api_key=…)
await upload_documents(provider, [
    {"@type": "Recipe", "url": "https://ex.com/r/laksa", "name": "Singapore Laksa",
     "description": "A coconut curry noodle soup."},
], site="recipes.example.com")
NLWeb GoodMem
url metadata["url"] — the item's identity, and what NLWeb dedupes on
site metadata["site"] — what every search filters on
raw_schema_object metadata["schema_json"]
(text to embed) the memory's content

The embedded text is not the raw JSON. Embedding a JSON blob buries the words a query would match under punctuation and key names, so the text is extracted from name, headline, description, articleBody and friends, and the untouched object is kept alongside for NLWeb to return verbatim.

upload_documents inspects every item of the batch response: the server returns HTTP 200 even when an item failed, so per-item success is the only signal. A partial failure raises GoodMemUploadError carrying the ids that did land.

Looking objects up by URL

from nlweb_goodmem import GoodMemObjectLookupProvider

lookup = GoodMemObjectLookupProvider(space_name="nlweb", base_url=…, api_key=…)
await lookup.get_by_id("https://ex.com/r/laksa")   # the full Schema.org object

Implements ObjectLookupProvider, so NLWeb can enrich a truncated search result with the complete object without a second datastore. NLWeb does that when an object_storage provider named default is configured:

object_storage:
  default:
    import_path: nlweb_goodmem
    class_name: GoodMemObjectLookupProvider
    base_url: https://localhost:8080
    api_key: gm_…
    space_name: nlweb

The id NLWeb passes here comes from a web request. It is the item's URL, not a GoodMem id, so it is not required to be a UUID: it only ever travels as an escaped value in the filter query parameter, never in the URL path.

Scores, and why there are none

A GoodMem vector score is a negative inner product — the best match is the lowest number — and a reranker score is a different scale that also goes negative. RetrievedItem has no score field, and inventing one would imply a comparability that does not hold. Results keep the server's ordering, which is authoritative, and are never re-sorted here.

Degraded retrieval

If the server reports a problem, whatever it did return is still returned and the statuses are logged. If it reports a problem and returns nothing, the result is an empty list plus a UserWarning — never an exception, because an exception here would take down an ask request that could still answer from another endpoint. Notices that carry no loss (FEATURE_DISABLED, LLM_CAPABILITY_INFERRED) are ignored; a status code this version does not know is reported as UNKNOWN rather than dropped.

Which NLWeb?

This targets nlweb-core (the pip-installable package with the config-driven provider architecture). The nlweb-ai/NLWeb reference implementation has a different interface (VectorDBClientInterface, returning list[list[str]]) and a hardcoded provider table, so a third-party package cannot register with it — that one needs an upstream PR.

Development

pip install -e ".[dev]"
ruff check src tests examples && mypy && pytest -m "not integration"

The offline suite replays NDJSON captured from a live GoodMem server (v1.0.320) through the real SDK decoders. The live suite needs a server:

GOODMEM_BASE_URL=… GOODMEM_API_KEY=… GOODMEM_EMBEDDER_ID=… \
  GOODMEM_RERANKER_ID=… GOODMEM_VERIFY_SSL=0 \
  pytest -m integration

GOODMEM_RERANKER_ID is optional — the reranker test skips without it. GOODMEM_VERIFY_SSL=0 is for a local server with a self-signed certificate.

There is no default credential anywhere in this repository.

License

MIT

Metadata

Release files for nlweb-goodmem 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nlweb-goodmem 0.2.1
File Size Uploaded
nlweb_goodmem-0.2.1.tar.gz 37.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nlweb-goodmem 0.2.1
File Interpreter ABI Platform
nlweb_goodmem-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 61.8 kB

Release files / nlweb_goodmem-0.2.1.tar.gz

Download URL nlweb_goodmem-0.2.1.tar.gz
Size 37.3 kB
Tags Source
SHA-256 checksum
How to use checksums
acd50d1ee9692e2b08b1fdc556337dd374e4247d2337b33a40122bcb912c93fb
BLAKE2b-256 checksum
How to use checksums
db44e5141e35e673355b914e990d3f3cbf86c95b6f1332ec459916ee80c7ce1f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / nlweb_goodmem-0.2.1-py3-none-any.whl

Download URL nlweb_goodmem-0.2.1-py3-none-any.whl
Size 24.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
de8b88158fabf19d312ccd24b30d25fa8cbd36b2b151749ef309f43e7f8f21bd
BLAKE2b-256 checksum
How to use checksums
803ebc61db411a8ea42bd8f9c19e725d486f1927b4ef20ba00d6c391690d669b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.2

2 release files

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page