Skip to main content

goodmem-deepeval

GoodMem retrieval for DeepEval.

Every retrieval is a DeepEval retriever span, and a retrieval turns directly into an LLMTestCase with retrieval_context populated — which is what ContextualPrecisionMetric, ContextualRecallMetric, ContextualRelevancyMetric and FaithfulnessMetric actually read.

pip install goodmem-deepeval

Retrieve and evaluate

from deepeval import evaluate
from deepeval.metrics import ContextualRelevancyMetric, ContextualRecallMetric
from goodmem_deepeval import GoodMemRetriever, GoodMemRetrievalHealthMetric

retriever = GoodMemRetriever(space_name="docs", limit=5)   # credentials from the environment

result = retriever.search("how do I rotate an API key?")
answer = my_llm(result["hits"])                            # your generation step
expected = "…"                                             # your reference answer

case = retriever.to_test_case(result, actual_output=answer, expected_output=expected)

evaluate([case], [
    ContextualRelevancyMetric(),
    ContextualRecallMetric(),
    GoodMemRetrievalHealthMetric(),
])

ContextualRecallMetric and ContextualPrecisionMetric compare against a reference answer, so they need expected_output; without it DeepEval's evaluate() raises MissingTestCaseParamsError. ContextualRelevancyMetric and FaithfulnessMetric do not.

search() is decorated with @observe(type="retriever") and reports top_k and the embedder through update_retriever_span, so the call appears as a retriever span with its query, latency and results.

If you only need the context, retrieve() returns the chunk text directly:

context = retriever.retrieve("how do I rotate an API key?")   # list[str]

What search() returns

Key What it is
query The query, as passed
hits chunk_id, chunk_text, memory_id, space_id, source, score, score_kind, metadata — in the server's order
score_kind "vector" or "reranker": what the server returned, not what was configured. Different scales; see below
statuses Statuses indicating a real problem, [] when clean
partial True when the server reported a real problem during this retrieval, with or without hits
abstract_reply The server-generated summary, only when llm_id is set
space_ids Which spaces were actually searched

A retrieval that reported a problem and returned nothing is an empty hits with partial: True and a warning — never an exception, and never indistinguishable from "no matches".

The health metric

from deepeval import evaluate
from deepeval.metrics import ContextualRecallMetric
from goodmem_deepeval import GoodMemRetrievalHealthMetric

# cases: LLMTestCases built with retriever.to_test_case(...), as above
evaluate(cases, [ContextualRecallMetric(), GoodMemRetrievalHealthMetric()])

Every other RAG metric scores the content that came back, so a broken reranker and a thin corpus look identical: both are poor recall. GoodMemRetrievalHealthMetric reads the diagnostics to_test_case puts in metadata["goodmem"] and scores the retrieval itself — 1.0 clean, 0.0 degraded, with the server's status codes in reason and score_breakdown. No LLM, so it costs nothing and never flakes. A test case with no GoodMem metadata is skipped rather than failed.

Scores

GoodMem returns two different things in the same field:

Range observed live Best match is
Vector score negative, e.g. -0.57 the lowest number
Reranker score can also go negative the highest number

So results keep the server's order and are never re-sorted; score_kind says which scale you have; and min_score is applied client-side, only to reranker scores. The server's relevance_threshold is never sent.

score_kind says what the server actually did, not what was configured. When reranker_id is set but the reranker fails, the server reports RERANKING_FAILED (and NOT_FOUND for a missing reranker) and still returns the vector-stage hits. Those hits are score_kind: "vector" with their raw vector scores, and min_score is not applied to them, so a reranker threshold cannot discard what the server returned; partial is True and statuses carries both codes. (0.2.1 labelled them "reranker", and live, with a missing reranker, min_score=0.0 returned none of the 3 hits the server sent.)

Even with a reranker the scale is model-dependent: on the same documents Voyage rerank-2.5 scored 0.27..0.93 and Jina jina-reranker-v3 scored -0.14..0.43. A min_score tuned for one empties the other, so when a threshold removes every hit the retriever warns and names the observed range. Calibrate min_score for the reranker you use; there is no default.

Filtering

retriever = GoodMemRetriever(space_name="docs", metadata_filter={"category": "billing"})
retriever.search("refunds", metadata_filter={"lang": "en", "archived": False})  # AND-ed per call

Each value is compared as its own type; the Python type picks the cast:

metadata_filter Sent to the server
{"category": "billing"} CAST(val('$.category') AS TEXT) = 'billing'
{"archived": False} CAST(val('$.archived') AS BOOLEAN) = false
{"year": 2026} CAST(val('$.year') AS NUMERIC) = 2026
{"score": 2.5} CAST(val('$.score') AS NUMERIC) = 2.5

Pass the type your metadata stores. Live (v1.0.320), 0.2.1 sent {"flag": True} as AS TEXT = 'True', which the server accepts with HTTP 200 and matches nothing, so a boolean filter looked like "nothing stored"; and {"n": 5.0} as '5.0', which misses a stored 5. AS BOOLEAN and AS NUMERIC match both. None and any other type (lists, dicts, bytes, NaN, infinity) raise ValueError before a request rather than being turned into text that silently matches nothing.

Text values are quoted for the GoodMem filter grammar (backslash escaping, verified against a live server; control characters refused). For anything more complex, pass an expression directly with filter=....

Attaching to a space by name

GoodMemRetriever(space_name="docs", embedder_id="…")                     # reuse or fail
GoodMemRetriever(space_name="docs", embedder_id="…", create_space=True)  # or create it

Attach-by-name is idempotent reuse: an existing space whose embedder matches is reused; a different embedder is an error, because retrieving across mismatched embedders returns plausible-looking nonsense. An ambiguous name is an error too.

Credentials

From GOODMEM_BASE_URL and GOODMEM_API_KEY, or as constructor keywords:

GoodMemRetriever(space_name="docs", base_url="https://goodmem.example.com",
                 api_key="gm_…")

verify_ssl defaults to on and exists for a local server with a self-signed certificate only; no example here turns it off.

They are deliberately not model fields, so they cannot reach a model_dump(), a repr() or a trace payload.

Development

pip install -e ".[dev]"
ruff check src tests examples && mypy && pytest -m "not integration"

CI runs the same three on Python 3.10–3.13, then builds the sdist and wheel, imports the wheel in a clean virtualenv, and fails if a GoodMem API key (gm_ followed by 20 or more lowercase letters or digits) is anywhere in the tree.

The offline suite replays NDJSON captured from a live GoodMem server (v1.0.320) through the real SDK decoders, so the wire format is never invented. The live suite needs a server and is skipped without one:

GOODMEM_BASE_URL=… GOODMEM_API_KEY=… GOODMEM_EMBEDDER_ID=… \
  GOODMEM_RERANKER_ID=… SSL_CERT_FILE=/path/to/local-ca.pem \
  pytest -m integration

GOODMEM_RERANKER_ID is optional — the working-reranker test skips without it. For a local server with a self-signed certificate, point SSL_CERT_FILE at its CA so TLS verification stays on (GOODMEM_VERIFY_SSL=0 turns it off). GOODMEM_E2E_SPACE_PREFIX names the temporary test space (default goodmem-deepeval-e2e-), so its owner is recognisable on a shared server.

There is no default credential anywhere in this repository.

License

MIT

Metadata

Release files for goodmem-deepeval 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for goodmem-deepeval 0.3.0
File Size Uploaded
goodmem_deepeval-0.3.0.tar.gz 38.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for goodmem-deepeval 0.3.0
File Interpreter ABI Platform
goodmem_deepeval-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 63.5 kB

Release files / goodmem_deepeval-0.3.0.tar.gz

Download URL goodmem_deepeval-0.3.0.tar.gz
Size 38.5 kB
Tags Source
SHA-256 checksum
How to use checksums
795831b76de56ae780efe6d1f6ce2e632724783bfc4a0f24f012cf1e63b1ff71
BLAKE2b-256 checksum
How to use checksums
2d0e07bbefbfb30735554c78b123a836958fd450b35771fc84f5d983ac61fd88
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / goodmem_deepeval-0.3.0-py3-none-any.whl

Download URL goodmem_deepeval-0.3.0-py3-none-any.whl
Size 25.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0724108bf921a987a9fdffd3da477c353c67a1ca875e3a014077af9b41074d77
BLAKE2b-256 checksum
How to use checksums
3741ab1a23be02a6daf6e05514399e380dffbbc9f04c4896cc8b3ac6138cf9c6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page