Skip to main content

PyVectorHound

Diagnose why your RAG retrieval is returning the wrong documents.

PyVectorHound is a component-level diagnostic engine for retrieval-augmented generation (RAG) pipelines. Point it at a set of search results (from your own pipeline, or from a live Qdrant/Chroma/Milvus/pgvector/Weaviate instance) and it isolates which stage is failing — embedding quality or vector search ranking — and gives you plain-English, ranked recommendations.

PyPI Python 3.8+ Tests License: Proprietary


What this actually does today

PyVectorHound does not run an embedding model, a reranker, or BM25 for you, and it is not a vector database. It's a diagnostic layer that sits on top of retrieval results you already have (or that it fetches from your vector DB) and tells you, with real computed metrics, what's wrong:

  • Embedding-space diagnostics (isotropy, coverage, distinctiveness) — computed by a Rust extension (pyvectorhound._core, built via PyO3) from the real per-document embeddings your database adapter's get_embeddings() returns.
  • Vector search accuracy (precision, recall, MRR) — computed against expected_docs you supply as ground truth.
  • LLM-as-judge faithfulness / contradiction checking — pass document_texts and an llm_judge_fn (no LLM client bundled, same pattern as embed_fn) and Diagnosis will flag retrieved documents that score well on embedding similarity but are actually irrelevant to or contradict the query — the failure mode pure distance metrics can't see.
  • Root cause + ranked recommendations — plain-English output combining the above.
  • Concurrent batch evaluation — Hound.diagnose_batch() runs many queries at once on a thread pool (with automatic retry/backoff on embed_fn/llm_judge_fn calls) instead of one at a time, for large-scale evaluation runs.
  • Model-agnostic quality thresholds — once you've tracked a few diagnoses with Hound.track_metric(), GOOD/MODERATE/WEAK status is computed relative to your own historical baseline instead of a fixed cutoff, so it doesn't drift when you switch embedding models.

BM25 (keyword search) and reranker diagnostics are reported as "UNKNOWN": PyVectorHound doesn't run a keyword-search index or a reranker itself, and Diagnosis doesn't yet accept external BM25/reranker scores as input, so rather than fabricate a number for a component it can't measure, it says so.

If a component doesn't have enough input to measure honestly (no adapter, fewer than 2 documents with embeddings, or no expected_docs), it's reported as "UNKNOWN" with an explanation of what to supply — never a made-up number.

PyVectorHound does not bundle an embedding model. If you want Hound.diagnose() to embed your query text for you, pass it an embed_fn (a thin wrapper around whatever you already use — OpenAI, Cohere, sentence-transformers, etc.). Without one, pass a precomputed query_embedding per call. It will not silently generate a random vector and pretend the resulting diagnosis means something.

New: advanced retrieval ranking (Rust core)

src/retrieval_ranking.rs adds a RetrievalRanker that combines BM25, semantic, recency, and diversity signals into a single multi-criteria ranking, plus cross-encoder-style reranking support. It's compiled into the native _core extension but not yet exposed as a Python-callable function — if you need it from Python today, treat it as in-progress internal infrastructure rather than a public API.

Not yet real (known limitations)

Being upfront about what's still a stub, rather than leaving it to look finished:

  • ModelComparison / Hound.compare_models() reports real, published cost/latency metadata for known models, but has no way to measure quality (F1/NDCG) on its own — pass quality_fn for real numbers, or it reports quality as unmeasured.
  • Hound.compare_metrics() and Hound.detect_drift() raise NotImplementedError; QualityScorer.trend_analysis() instead returns a dict with "direction": "unknown" and an explanation. None of the three has a historical data store to compute a real trend from. Use Hound.track_metric() + Hound.get_trend_report() (backed by the real, tested TrendAnalyzer) instead.
  • As of v1.3.1, PyPI only carries a macOS arm64 / CPython 3.11 wheel and no source distribution. On any other interpreter or OS, pip install pyvectorhound does not build the current version from source — it silently falls back to the last release that does have an sdist (currently 1.3.0, three releases behind), with no warning that you got an old version. If you need the current release outside macOS arm64/CPython 3.11, install straight from the repo instead, which does build the latest source correctly (needs a Rust toolchain — maturin/pip handle the build, but cargo must be available): pip install git+https://github.com/Mullassery/PyVectorHound.git.
  • GitHub Actions CI (the badge above) is currently red on every job across recent pushes to main — not because of failing tests, but because .github/workflows/ci.yml's dtolnay/rust-toolchain@v1 step is missing its required toolchain input, so every job fails in the setup step before any code runs. The local test suite itself passes (159/159 as of this writing, run via pytest tests/ -v with the Rust extension built).
  • OpenTelemetry / LangChain / LlamaIndex / MCP integrations, the CLI, and the REST server exist and have passing tests but have seen far less real-world use than the core Hound/Diagnosis path above.

Installation

pip install pyvectorhound

Optional vector database clients (only install the one(s) you use):

pip install pyvectorhound[qdrant]     # Qdrant
pip install pyvectorhound[chroma]     # Chroma
pip install pyvectorhound[milvus]     # Milvus
pip install pyvectorhound[weaviate]   # Weaviate
pip install pyvectorhound[pgvector]   # PostgreSQL + pgvector

Requires Python 3.8+.


Quick start: diagnose results you already have

This is the fastest way to try it — no live database or embedding model needed. Diagnosis fetches per-document embeddings for you via a small adapter object (anything with a get_embeddings(doc_ids) -> dict method); without one, the embedding component honestly reports "UNKNOWN" instead of a fabricated score.

from pyvectorhound import Diagnosis

class InMemoryAdapter:
    """Anything with get_embeddings(doc_ids) works -- swap in your own
    QdrantAdapter/ChromaAdapter/etc., or a wrapper around your pipeline."""
    def __init__(self, embeddings_by_id):
        self._embeddings_by_id = embeddings_by_id

    def get_embeddings(self, doc_ids):
        return {d: self._embeddings_by_id[d] for d in doc_ids if d in self._embeddings_by_id}

results = [
    {"id": "pricing.pdf", "score": 0.91},
    {"id": "onboarding.md", "score": 0.84},
    {"id": "faq.md", "score": 0.79},
]

diagnosis = Diagnosis(
    query="What's your return policy?",
    results=results,
    expected_docs=["returns.pdf", "policy.md"],  # ground truth
    adapter=InMemoryAdapter(my_document_embeddings),
)
diagnosis.analyze()

print(diagnosis.root_cause())
for rec in diagnosis.recommendations():
    print(f"[{rec['priority']}] {rec['action']}")

print(diagnosis.hunt())  # full plain-English report

A runnable version (with synthetic embeddings so it works with no setup) is in examples/retrieval_debug.py.

Quick start: diagnose against a live vector database

from pyvectorhound import Hound

hound = Hound(
    db="qdrant",                      # qdrant | chroma | milvus | weaviate | postgres
    endpoint="localhost:6333",
    index_name="documents",
    # PyVectorHound doesn't ship an embedding model -- wrap whatever you use:
    embed_fn=lambda text: my_embedding_client.embed(text),
)

diagnosis = hound.diagnose(
    query="What's your return policy?",
    expected_docs=["returns.pdf", "policy.md"],
    top_k=5,
)
print(diagnosis.hunt())

Hound connects lazily — constructing it doesn't require a live server, only calling diagnose() (or another querying method) does. diagnose() already passes self.adapter into Diagnosis, so embedding-space diagnostics work out of the box against your real database.


Diagnostics it runs

Component What it measures Requires
Embedding Isotropy, coverage, distinctiveness of the retrieved documents' real embeddings An adapter with get_embeddings(), and ≥2 retrieved documents
Vector search Precision, recall, MRR expected_docs (ground truth)
Faithfulness LLM-judge contradiction/relevance check document_texts + llm_judge_fn
BM25 (keyword) Not implemented — reports UNKNOWN n/a
Reranker Not implemented — reports UNKNOWN n/a

Every measured component is computed for real from the input you give it; nothing is guessed when the input isn't there.


Faithfulness checking and batch evaluation

from pyvectorhound import Hound

def judge(query: str, doc_texts: list[str]) -> dict:
    # Wrap whatever LLM client you already use -- PyVectorHound doesn't
    # bundle one. Must return at least a "contradiction_score" (0.0-1.0,
    # lower is more faithful) and/or "faithful"/"contradicted_count".
    response = my_llm_client.judge_faithfulness(query, doc_texts)
    return {
        "contradiction_score": response.score,
        "contradicted_count": response.contradicted,
        "reasoning": response.explanation,
    }

hound = Hound(db="qdrant", embed_fn=my_embed_fn, llm_judge_fn=judge)

diagnosis = hound.diagnose(
    query="What's your return policy?",
    document_texts={"returns.pdf": "...", "policy.md": "..."},  # doc_id -> text
)
print(diagnosis.metrics()["faithfulness"])

# Evaluate many queries concurrently instead of one at a time:
diagnoses = hound.diagnose_batch(
    queries=["query 1", "query 2", "query 3"],
    document_texts=[{"a": "..."}, None, {"b": "..."}],  # per-query, optional
    max_workers=8,
)

Track a few diagnoses over time and quality-status classification switches from a fixed cutoff to your own historical baseline automatically:

hound.track_metric("vector_search_precision", diagnosis.metrics()["vector_search"]["precision"])
# After ~5+ tracked points, later diagnose() calls classify status
# (GOOD/MODERATE/WEAK) relative to that baseline instead of a fixed number.

Other tools

  • hound.quality_scorer() — QualityScorer for scoring an embedding's validity, and (given corpus neighbors via the adapter) real isotropy/coverage/distinctiveness against the corpus.
  • hound.benchmark() — PerformanceBenchmark for latency percentiles and database/embedding-model comparisons.
  • hound.analyze_trends() — TrendAnalyzer for tracking metrics over time and detecting drift, regressions, and anomalies from real tracked values.
  • hound.tracer() / hound.replayer() — capture a retrieval pipeline run and replay it under different configurations to compare recall/latency.

See examples/ for runnable scripts, and docs/ARCHITECTURE.md / docs/GUIDE.md for more detail.


Development

git clone https://github.com/Mullassery/PyVectorHound.git
cd PyVectorHound
pip install maturin
maturin develop --release   # builds the Rust extension in place
pip install -e ".[dev]"
pytest tests/ -v

License

Proprietary License — free to use with explicit attribution. See LICENSE.

Release files for pyvectorhound 1.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyvectorhound 1.4.0
File Size Uploaded
pyvectorhound-1.4.0.tar.gz 178.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyvectorhound 1.4.0
File Interpreter ABI Platform
pyvectorhound-1.4.0-cp39-cp39-macosx_11_0_arm64.whl CPython 3.9 CPython 3.9 macOS 11.0+ ARM64 Details

Total release size: 485.5 kB

Release files / pyvectorhound-1.4.0.tar.gz

Download URL pyvectorhound-1.4.0.tar.gz
Size 178.4 kB
Tags Source
SHA-256 checksum
How to use checksums
e6dc1f12c16051d3827ae6bb06acfd41cab08473e84c630faaa152c46aee1581
BLAKE2b-256 checksum
How to use checksums
f67151b43959fa5fa41e4819821ff5d3a2d90fc17b478e98ca873bcd0826c1ff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / pyvectorhound-1.4.0-cp39-cp39-macosx_11_0_arm64.whl

Download URL pyvectorhound-1.4.0-cp39-cp39-macosx_11_0_arm64.whl
Size 307.1 kB
Tags CPython 3.9 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
f31d15d87ebbc95127cd97eaab35b08c2b0527ec81a079c631e52160ecf4c130
BLAKE2b-256 checksum
How to use checksums
006964935bc244b019fe7991369f11363c1fbebd0e1b6032b2abd001c1b8852f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

1.5.0

2 release files

This release

1.4.0 This release

2 release files

1.3.3

1 release file

1.3.2

1 release file

1.3.1

1 release file

1.3.0

3 release files

1.2.1

1 release file

1.2.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page