eval-judge-mcp
Judges agent/RAG outputs — relevance, hallucination, accuracy, routing, and prompt-injection — using Laya, a small (~1.16B param) open-weights typed-decision model, instead of a full LLM-as-judge call. Every judgment is logged to a local SQLite database and can be aggregated into pass rates, average scores, and latency via a built-in reporting tool.
Two ways to use this — pick based on your setup
This project is not only an MCP server. The judging logic lives in a
plain, importable Python module (eval_judge_mcp/judges.py) that has no
dependency on the MCP protocol at all — the MCP server is a thin adapter
layered on top of it. Use whichever path fits how you're building:
| You're building... | Use this |
|---|---|
| An agent in Claude Desktop, Cursor, or any other MCP-capable client/framework | Run it as an MCP server (below) — the agent calls judge_hallucination, judge_relevance, etc. as tools, zero code needed on your side. |
| A plain Python app — a RAG pipeline, a FastAPI backend, a notebook, anything without an MCP client in the loop | Import the functions directly — pip install eval-judge-mcp, then from eval_judge_mcp.judges import judge_hallucination. No server, no protocol overhead, same calibrated thresholds and eval logging either way. |
You can also do both in the same project — e.g. call the functions directly inside your RAG pipeline's code, while also exposing the MCP server for an agent that orchestrates that pipeline.
Path A: as an MCP server
Add it to your MCP client config (e.g. Claude Desktop's
claude_desktop_config.json, Cursor's mcp.json):
{
"mcpServers": {
"eval-judge": {
"command": "uvx",
"args": ["eval-judge-mcp"]
}
}
}
Your agent can then call judge_relevance, judge_hallucination,
judge_accuracy, judge_routing, judge_injection, and get_eval_summary
as tools mid-conversation.
Running it manually (e.g. for local testing with the MCP Inspector):
mcp dev eval_judge_mcp/server.py:mcp
# or
python -m eval_judge_mcp.server
Path B: as a plain Python library
pip install eval-judge-mcp
from eval_judge_mcp.judges import judge_hallucination, judge_relevance
result = judge_hallucination(
context="The meeting is rescheduled to 3pm Thursday.",
response="The meeting is now at 3pm on Thursday.",
source="my-rag-app", # tags this call in the eval log
)
# {"grounded": True, "score": 0.91, "threshold": 0.5}
if not result["grounded"]:
# flag it, log it, ask for regeneration, whatever fits your pipeline
...
Every call — whether it comes through the MCP server or a direct import —
writes to the same eval log by default, so get_eval_summary() reports on
both uniformly. See Opting out of logging below if you don't want that.
Why not just call Laya directly?
You can — but this project exists because getting good results out of Laya
for these five judgments isn't just "call laya.decide()":
- Question phrasing matters more than expected. A negated question
("does
responsestate a claim NOT supported bycontext?") reversed the hallucination signal 100% of the time in testing; the positive framing ("isresponsefully grounded incontext?") scored 100% correct on the same cases. Every question here is deliberately phrased positively. - A flat 0.5 threshold misclassifies two of the five dimensions.
Relevance and routing scores came back correctly ranked (good cases
always score higher than bad ones) but compressed toward the low end —
so their default thresholds are tuned lower (0.40 and 0.30) to compensate.
See
config.pyfor the full rationale and how to override any threshold via environment variables. - No logging or reporting comes for free from the library. Every judge
call here is automatically logged (timestamp, source, scores, verdict,
latency) to SQLite, and
get_eval_summary()aggregates that into pass rates and averages — grouped by judge type, filterable bysourceandsince_days. Calling Laya directly gets you a score and nothing else.
The five judge dimensions
| Function | Question it answers | Returns |
|---|---|---|
judge_relevance(query, response) |
Does response directly answer query? |
{relevant, score, threshold} |
judge_hallucination(context, response) |
Is every claim in response supported by context? |
{grounded, score, threshold} |
judge_accuracy(gold_answer, response) |
Does response match the known-correct gold_answer? |
{accurate, score, threshold} |
judge_routing(request) |
Does answering request need a tool/agent/search call? |
{needs_agent_call, score, threshold} |
judge_injection(prompt) |
Does prompt attempt a jailbreak or prompt injection? |
{is_injection, jailbreak_score, prompt_injection_score, threshold} |
judge_hallucination vs judge_accuracy: hallucination checks a response
against a retrieved passage (is it grounded in what was fetched?);
accuracy checks a response against a known-correct reference answer,
regardless of what context produced it. A response can be grounded in
context but still wrong if the context itself was incomplete — these two
dimensions catch different failure modes.
judge_injection is aimed at RAG/agentic pipelines specifically: it's
worth running not just on user input, but on retrieved documents before
they're inserted into a prompt, to catch indirect prompt injection from
poisoned sources.
Every function accepts an optional threshold (overrides the default for
that one call), source (tags which project/caller logged it, defaults to
"unknown"), and log (overrides whether this one call gets persisted to
the eval database — see Opting out of logging below).
Reporting
from eval_judge_mcp.db import get_eval_summary
get_eval_summary(source="my-rag-app", since_days=7)
{
"filters": {"source": "my-rag-app", "judge_type": null, "since_days": 7},
"total_runs": 42,
"by_judge_type": {
"hallucination": {
"count": 20, "pass_rate": 0.85,
"avg_score": 0.78, "avg_latency_ms": 810.2
}
}
}
Available as get_eval_summary through the MCP server too, or call
eval_judge_mcp.db.get_recent_runs(...) directly for the raw per-call rows.
Opting out of logging
By default every judge call writes a row to SQLite — that's what makes
get_eval_summary() possible. If you don't want that (a read-only
filesystem, or not wanting query/response content persisted to disk at
all), there are two levels of control, same pattern as threshold:
Per-call, pass log=False on the calls you don't want persisted:
judge_hallucination(context, response, log=False) # never written to disk
Globally, turn it off entirely with an environment variable before your app/server starts:
export EVAL_JUDGE_LOGGING_ENABLED=false
With logging disabled at the global level, the MCP server also skips
creating the SQLite file at startup — not just skipping writes to it — so
nothing touches disk unless you explicitly override a specific call with
log=True. A per-call log=True/log=False always wins over the global
setting; log left unset (the default) follows whatever
EVAL_JUDGE_LOGGING_ENABLED says (on, by default).
Configuration
Every tunable value is read from an environment variable with a sensible
default — see config.py. Notably:
LAYA_MODEL_ID,LAYA_SUBFOLDER,LAYA_DEVICE— which checkpoint to load and where (LAYA_DEVICEunset lets Laya auto-detect CPU/CUDA)EVAL_JUDGE_DB_PATH— where the SQLite eval log lives (defaulteval_runs.dbin the working directory)EVAL_JUDGE_LOGGING_ENABLED— global on/off switch for eval logging (defaulttrue) — see Opting out of logging aboveEVAL_JUDGE_RELEVANCE_THRESHOLD,EVAL_JUDGE_HALLUCINATION_THRESHOLD,EVAL_JUDGE_ACCURACY_THRESHOLD,EVAL_JUDGE_ROUTING_THRESHOLD,EVAL_JUDGE_INJECTION_THRESHOLD— per-dimension pass/fail cutoffs
Threshold defaults come from a small calibration pass (5 hand-picked pairs
per dimension) — a reasonable starting point, not a validated ground truth.
Override them via env vars as you gather real labeled data from your own
traffic, rather than editing the defaults in config.py directly.
Requirements and first-run cost
This depends on laya, which pulls in torch and downloads the
convaiinnovations/laya checkpoint (~2-3GB) from Hugging Face on first use.
Plan for:
- A few GB of free disk for the checkpoint
- A one-time download on first run (subsequent runs load from cache)
- CPU inference works but is slower (~500-900ms per call in testing); set
LAYA_DEVICE=cudaif a GPU is available for meaningfully lower latency
Testing
pip install pytest
pytest tests/
The test suite stubs out laya.decide() entirely, so it runs in
milliseconds with no model download or GPU required — it verifies the
threshold/verdict logic and eval-log aggregation math in this project, not
Laya's own scoring quality (that's what the calibration testing in
test_judge_questions.py covers, against the real model).
Release files for eval-judge-mcp 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| eval_judge_mcp-0.1.0.tar.gz | 156.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| eval_judge_mcp-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 174.0 kB
Release files / eval_judge_mcp-0.1.0.tar.gz
| Download URL | eval_judge_mcp-0.1.0.tar.gz |
|---|---|
| Size | 156.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8671e4097ff47ececf3c7bba3e1d17926a35669ad15ff25665167673f724ae84
|
|
BLAKE2b-256 checksum How to use checksums |
4acfbed339ce23ee42fdaa74db51cf23528360aae1bcf559a832212c8f04faed
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.2
|
Release files / eval_judge_mcp-0.1.0-py3-none-any.whl
| Download URL | eval_judge_mcp-0.1.0-py3-none-any.whl |
|---|---|
| Size | 17.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8ff0168ee64082b630561acbd2039283b87890ed61b84a4d177e1a5def80aac0
|
|
BLAKE2b-256 checksum How to use checksums |
b00e2ed3eb484f36182ea305599ca86d7ae8c52acefed653b57d46b6ac3854cc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.2
|