goodmem-deepeval
GoodMem retrieval for DeepEval.
Every retrieval is a DeepEval retriever span, and a retrieval turns
directly into an LLMTestCase with retrieval_context populated — which is
what ContextualPrecisionMetric, ContextualRecallMetric,
ContextualRelevancyMetric and FaithfulnessMetric actually read.
pip install goodmem-deepeval
Retrieve and evaluate
from deepeval import evaluate
from deepeval.metrics import ContextualRelevancyMetric, ContextualRecallMetric
from goodmem_deepeval import GoodMemRetriever, GoodMemRetrievalHealthMetric
retriever = GoodMemRetriever(space_name="docs", limit=5) # credentials from the environment
result = retriever.search("how do I rotate an API key?")
answer = my_llm(result["hits"]) # your generation step
expected = "…" # your reference answer
case = retriever.to_test_case(result, actual_output=answer, expected_output=expected)
evaluate([case], [
ContextualRelevancyMetric(),
ContextualRecallMetric(),
GoodMemRetrievalHealthMetric(),
])
ContextualRecallMetric and ContextualPrecisionMetric compare against a
reference answer, so they need expected_output; without it DeepEval's
evaluate() raises MissingTestCaseParamsError.
ContextualRelevancyMetric and FaithfulnessMetric do not.
search() is decorated with @observe(type="retriever") and reports top_k
and the embedder through update_retriever_span, so the call appears as a
retriever span with its query, latency and results.
If you only need the context, retrieve() returns the chunk text directly:
context = retriever.retrieve("how do I rotate an API key?") # list[str]
What search() returns
| Key | What it is |
|---|---|
query |
The query, as passed |
hits |
chunk_id, chunk_text, memory_id, space_id, source, score, score_kind, metadata — in the server's order |
score_kind |
"vector" or "reranker": what the server returned, not what was configured. Different scales; see below |
statuses |
Statuses indicating a real problem, [] when clean |
partial |
True when the server reported a real problem during this retrieval, with or without hits |
abstract_reply |
The server-generated summary, only when llm_id is set |
space_ids |
Which spaces were actually searched |
A retrieval that reported a problem and returned nothing is an empty hits
with partial: True and a warning — never an exception, and never
indistinguishable from "no matches".
The health metric
from deepeval import evaluate
from deepeval.metrics import ContextualRecallMetric
from goodmem_deepeval import GoodMemRetrievalHealthMetric
# cases: LLMTestCases built with retriever.to_test_case(...), as above
evaluate(cases, [ContextualRecallMetric(), GoodMemRetrievalHealthMetric()])
Every other RAG metric scores the content that came back, so a broken
reranker and a thin corpus look identical: both are poor recall.
GoodMemRetrievalHealthMetric reads the diagnostics to_test_case puts in
metadata["goodmem"] and scores the retrieval itself — 1.0 clean, 0.0
degraded, with the server's status codes in reason and score_breakdown.
No LLM, so it costs nothing and never flakes. A test case with no GoodMem
metadata is skipped rather than failed.
Scores
GoodMem returns two different things in the same field:
| Range observed live | Best match is | |
|---|---|---|
| Vector score | negative, e.g. -0.57 |
the lowest number |
| Reranker score | can also go negative | the highest number |
So results keep the server's order and are never re-sorted; score_kind
says which scale you have; and min_score is applied client-side, only to
reranker scores. The server's relevance_threshold is never sent.
score_kind says what the server actually did, not what was configured. When
reranker_id is set but the reranker fails, the server reports
RERANKING_FAILED (and NOT_FOUND for a missing reranker) and still returns
the vector-stage hits. Those hits are score_kind: "vector" with their raw
vector scores, and min_score is not applied to them, so a reranker threshold
cannot discard what the server returned; partial is True and statuses
carries both codes. (0.2.1 labelled them "reranker", and live, with a
missing reranker, min_score=0.0 returned none of the 3 hits the server sent.)
Even with a reranker the scale is model-dependent: on the same documents
Voyage rerank-2.5 scored 0.27..0.93 and Jina jina-reranker-v3 scored
-0.14..0.43. A min_score tuned for one empties the other, so when a
threshold removes every hit the retriever warns and names the observed range.
Calibrate min_score for the reranker you use; there is no default.
Filtering
retriever = GoodMemRetriever(space_name="docs", metadata_filter={"category": "billing"})
retriever.search("refunds", metadata_filter={"lang": "en", "archived": False}) # AND-ed per call
Each value is compared as its own type; the Python type picks the cast:
metadata_filter |
Sent to the server |
|---|---|
{"category": "billing"} |
CAST(val('$.category') AS TEXT) = 'billing' |
{"archived": False} |
CAST(val('$.archived') AS BOOLEAN) = false |
{"year": 2026} |
CAST(val('$.year') AS NUMERIC) = 2026 |
{"score": 2.5} |
CAST(val('$.score') AS NUMERIC) = 2.5 |
Pass the type your metadata stores. Live (v1.0.320), 0.2.1 sent
{"flag": True} as AS TEXT = 'True', which the server accepts with HTTP 200
and matches nothing, so a boolean filter looked like "nothing stored"; and
{"n": 5.0} as '5.0', which misses a stored 5. AS BOOLEAN and
AS NUMERIC match both. None and any other type (lists, dicts, bytes, NaN,
infinity) raise ValueError before a request rather than being turned into
text that silently matches nothing.
Text values are quoted for the GoodMem filter grammar (backslash escaping,
verified against a live server; control characters refused). For anything
more complex, pass an expression directly with filter=....
Attaching to a space by name
GoodMemRetriever(space_name="docs", embedder_id="…") # reuse or fail
GoodMemRetriever(space_name="docs", embedder_id="…", create_space=True) # or create it
Attach-by-name is idempotent reuse: an existing space whose embedder matches is reused; a different embedder is an error, because retrieving across mismatched embedders returns plausible-looking nonsense. An ambiguous name is an error too.
Credentials
From GOODMEM_BASE_URL and GOODMEM_API_KEY, or as constructor keywords:
GoodMemRetriever(space_name="docs", base_url="https://goodmem.example.com",
api_key="gm_…")
verify_ssl defaults to on and exists for a local server with a self-signed
certificate only; no example here turns it off.
They are deliberately not model fields, so they cannot reach a
model_dump(), a repr() or a trace payload.
Development
pip install -e ".[dev]"
ruff check src tests examples && mypy && pytest -m "not integration"
CI runs the same three on Python 3.10–3.13, then builds the sdist and wheel,
imports the wheel in a clean virtualenv, and fails if a GoodMem API key
(gm_ followed by 20 or more lowercase letters or digits) is anywhere in the
tree.
The offline suite replays NDJSON captured from a live GoodMem server (v1.0.320) through the real SDK decoders, so the wire format is never invented. The live suite needs a server and is skipped without one:
GOODMEM_BASE_URL=… GOODMEM_API_KEY=… GOODMEM_EMBEDDER_ID=… \
GOODMEM_RERANKER_ID=… SSL_CERT_FILE=/path/to/local-ca.pem \
pytest -m integration
GOODMEM_RERANKER_ID is optional — the working-reranker test skips without
it. For a local server with a self-signed certificate, point SSL_CERT_FILE
at its CA so TLS verification stays on (GOODMEM_VERIFY_SSL=0 turns it off).
GOODMEM_E2E_SPACE_PREFIX names the temporary test space (default
goodmem-deepeval-e2e-), so its owner is recognisable on a shared server.
There is no default credential anywhere in this repository.
License
MIT
Metadata
Release files for goodmem-deepeval 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| goodmem_deepeval-0.3.0.tar.gz | 38.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| goodmem_deepeval-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 63.5 kB
Release files / goodmem_deepeval-0.3.0.tar.gz
| Download URL | goodmem_deepeval-0.3.0.tar.gz |
|---|---|
| Size | 38.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
795831b76de56ae780efe6d1f6ce2e632724783bfc4a0f24f012cf1e63b1ff71
|
|
BLAKE2b-256 checksum How to use checksums |
2d0e07bbefbfb30735554c78b123a836958fd450b35771fc84f5d983ac61fd88
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency logRelease files / goodmem_deepeval-0.3.0-py3-none-any.whl
| Download URL | goodmem_deepeval-0.3.0-py3-none-any.whl |
|---|---|
| Size | 25.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0724108bf921a987a9fdffd3da477c353c67a1ca875e3a014077af9b41074d77
|
|
BLAKE2b-256 checksum How to use checksums |
3741ab1a23be02a6daf6e05514399e380dffbbc9f04c4896cc8b3ac6138cf9c6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency log