Skip to main content

numguard

smithery badge

MCP registry identity — mcp-name: io.github.ipezygj/numguard

The verification layer for the agent economy — an agent-callable primitive that checks a number before it gets asserted, and hands back a signed receipt proving it was checked.

Agents now produce an explosion of numbers: eval scores, A/B results, "the agent improved 12%", benchmark rankings, backtest Sharpes. The scarce resource isn't the number — it's trust in the number. numguard is the tool an agent calls mid-task to ask "does this survive a second look?", and to attach a portable, tamper-evident receipt so the answer travels with the claim.

Built on evalgate for the shared eval statistics; adds the pieces agents specifically need — a Deflated Sharpe Ratio for backtests, judge calibration, signed receipts, and metering an agent can actually pay (prepaid credits + x402 pay-per-call). Exposed as an MCP server, so any agent can call it.

New here?How to verify a backtest is real (Deflated Sharpe in Python): the practical guide to catching an overfit or leaking backtest, with runnable code. New in 0.2.0: fdr_hurdle — no universal "t > 3"; derive the hurdle your own search history implies at your false-discovery-rate target (Harvey & Liu, JF 2020). See it workproof gallery: 8 real numbers run through the real checks, 3 survive and 5 are flagged, each with a receipt you can verify offline. Don't trust it — verify it. Came here for the statistics, not the plumbing? Start with numguard/fdr.py — the data-driven t-stat hurdle of Harvey & Liu, False (and Missed) Discoveries in Financial Economics, JF 2020: demean the trial panel, resample the time index with the same draws for every trial so the cross-trial correlation survives, and take the smallest hurdle whose estimated FDR meets your target. Its docstring states plainly what it is not — a single-bootstrap core, all trials treated as null when counting expected false discoveries, and the optional outer bootstrap reporting sampling variability of the hurdle rather than the paper's double-bootstrap p-value calibration. Tests: tests/test_fdr.py — the one worth a minute checks the estimator against an analytic value it was never told, E[#null ≥ h] = m·2(1−Φ(h)). The Deflated Sharpe Ratio lives in numguard/backtest.py. Pure math + seeded random; no numpy, no scipy.

Wire it into an agent in one lineINTEGRATE.md: the local reflex, an MCP config, and LangChain / CrewAI tool wrappers.


The tools

MCP tool What an agent asks it
verify_backtest Is this strategy's Sharpe real, or the luckiest of the many I tried? (Deflated Sharpe Ratio)
verify_backtest_series Run the full integrity battery on my actual returns — look-ahead, autocorrelation, regime, tail, overfitting.
verify_fdr_hurdle No universal t>3 — what t-stat hurdle does MY OWN search history imply at MY false-discovery-rate target? (Harvey & Liu 2020; pass the whole trial panel)
verify_subset_win Does "we lead on subset X" survive correcting for how many subsets I tested?
verify_model_gap Is the gap between these two models bigger than the test set can resolve?
verify_judge_bias Is my judge's preference real, or just longer / first / same-family?
calibrate_judge Is the LLM judge I trust actually calibrated against ground truth?
audit_leaderboard Is #1 on this leaderboard statistically real? (rank confidence intervals)
triage (start here) I don't know which check I need — here's what I'm about to do or assert, route me. (front door across numguard + agent-guard + evalgate, free)
verify_execution Don't trust my reported Sharpe — RE-DERIVE it from my positions on committed price data, and catch a number those decisions don't produce.
reconcile_backtest Did my backtest's claimed Sharpe survive contact with LIVE returns? (HELD / DECAYED / BROKEN)
open_commitment / report_returns Hold my strategy accountable over time — stream live returns, tell me when the edge breaks. (O(1)/obs)
open_precommitment / report_precommit Prove my live claim wasn't cherry-picked after the fact — pre-register it BEFORE the outcome; report on a hash-chained, tamper-evident timeline anyone can audit free (verify_chain).
issue_receipt / commitment_receipt Give me a signed, portable proof this number / track record was checked.
verify_receipt Was the number this other agent handed me actually checked, and by whom? (free, issuer-agnostic)
scan_for_receipts A peer just sent me a message — find and verify any receipt inside it before I act. (free — the receiver half of the loop)
receipt_spec / why / pricing / balance the open receipt standard · what numguard does that nothing else does · prices · balance

On-chain and agent-verification tools — the same discipline applied to things that live on a chain rather than in a spreadsheet. Listed because a tool an agent cannot find is a tool it cannot call.

tool the question it answers
verify_agent This wallet claims a track record — fetch its own public on-chain trades and re-derive the result.
audit_addresses Run that same verdict across an explicit list of addresses. (only the addresses given)
verify_vault Re-derive a vault's APY from its own Deposit/Withdraw events, instead of quoting its page.
verify_backing Re-derive backing = reserves held / token supply, from the chain.
verify_guard_trace Recompute a behavioural-guard verdict over an agent's action trace, and sign it.
anchor_receipt / attest_onchain Put a receipt's digest on Base — immutable, timestamped, publicly checkable (EAS attestation).
check_attestation Look up a numguard credential on-chain. (free, no key, no gas)
erc8004_feedback Build the ERC-8004 giveFeedback call from a verdict, so reputation carries the evidence. (free)
get_precommit / commitment_status The immutable registration entry, and the current HELD / DECAYED / BROKEN verdict. (free)

What sets it apart (why): computing the number yourself, or a lesser checker, stops at "is it significant?" numguard also holds it accountable to live reality over time, signs a portable tamper-evident proof, and lets anyone verify any proof for free — the trust layer, not just a calculator.

For agent traders: the Deflated Sharpe Ratio

The number that kills a backtest is the same one that kills a benchmark score: you tried many, and you reported the best. In finance the rigorous correction is the Deflated Sharpe Ratio (Bailey & López de Prado) — given how many variants you tested, what Sharpe would the luckiest zero-skill strategy have shown, and do you beat it after adjusting for sample length and non-normal returns?

from numguard import deflated_sharpe

deflated_sharpe(sr=0.12, T=250, n_trials=100)
# SR=0.120 over T=250, 100 trials tested; deflation bar=0.160; DSR=0.263
# -> does NOT survive deflation.  (PSR-vs-0=0.970 — it LOOKS significant on a single test.)

deflated_sharpe(sr=0.15, T=1000, n_trials=1)
# DSR=1.000 -> SURVIVES. A real edge over a long sample.

The contrast is the whole point: a single-test probability of 0.97 ("significant!") collapses to a deflated 0.26 ("noise") once you account for the 100 strategies that were tried. An agent optimizing over strategies should call this before it trusts — or publishes — a backtest.

The full integrity battery — what a Deflated Sharpe still misses

DSR catches best-of-N. It does not catch same-bar look-ahead, autocorrelation inflating the Sharpe, regime dependence, tail fantasy, or one-lucky-epoch fragility. verify_backtest_series runs the whole battery on the actual returns series and returns a risk level (none/medium/high/critical) plus the checks that flagged:

check catches
leakage same-bar look-ahead (position "predicts" the bar it's in) — critical
pbo overfitting beyond n_trials (Prob. of Backtest Overfitting) — critical
hac_sharpe autocorrelation / stale marks inflating the Sharpe (Newey–West)
regime_stability cherry-picked window (per-block Sharpe + CUSUM break)
bootstrap_stability edge lives in one epoch (block-bootstrap Sharpe CI)
drawdown tail/smoothing fantasy (Calmar / CVaR / expected-vs-realized max-DD)
permutation, conditional_hetero, cost_capacity, bh_fdr order structure, vol clustering, fill realism, multiple testing

The tell (python examples/catch_a_fake_backtest.py): a look-ahead strategy shows an annualised Sharpe of +20 and a Deflated Sharpe that survives — yet the battery flags it critical on leakage (same-bar corr 0.79 vs next-bar 0.05). The DSR waves the fiction through; the battery does not.

verify_backtest_series(api_key="…", returns=[...], positions=[...], asset_returns=[...])
# {"risk": "critical", "survives": false, "flags": ["leakage", ...], "checks": {...}}

Signed receipts (the part that compounds)

from numguard import verify_claim, issue_receipt, verify_receipt, keypair
priv, pub = keypair()
result  = verify_claim("backtest", sr=0.12, T=250, n_trials=100)
receipt = issue_receipt(result, priv, pub)     # Ed25519-signed
verify_receipt(receipt)                          # True — anyone can verify with the public key alone

Attach the receipt to your output. A downstream agent (or human) can confirm — without your keys — that the claim and its verdict weren't altered and that numguard issued them. As receipts circulate, "a number without a receipt" starts to read like "a number nobody checked."

Buying is easy for an agent

Two rails, both built so an agent can decide and pay in-loop, no human clicking:

  1. Prepaid credits + API key — a human tops up once; the agent spends per call. Generous free tier (25 calls/key) so the agent feels the value first, then a machine-readable price list. Insufficient balance returns a structured payment_required, not an error.
  2. x402 pay-per-call — the agent hits a tool, gets an HTTP-402 with a machine-readable price + pay-to address, pays USDC from its wallet, retries with proof, gets the result. The protocol layer is here; settlement is pluggable (inject a facilitator/RPC verifier for production).
from numguard import x402
x402.require_payment("verify_backtest", price_usd=0.03, pay_to="0x…")
# -> {"status": 402, "accepts": [{"scheme":"exact","network":"base","asset":"USDC", ...}]}

Run the MCP server

pip install numguard
python -m numguard.mcp_server        # stdio MCP server; point your agent/host at it

Then an agent calls e.g. verify_backtest(api_key="…", sr=0.12, T=250, n_trials=100) and gets a verdict it can quote and a receipt it can attach.

Deploy it (hosted, paid, discoverable)

1. Host the paid HTTP API (x402 per-call):

docker build -t numguard . && docker run -p 8080:8080 \
  -e NUMGUARD_PAYTO=0xYOURWALLET \
  -e NUMGUARD_FACILITATOR_URL=https://your-x402-facilitator \
  numguard

Or one-click on Render: New → Blueprint → this repo (render.yaml included); set NUMGUARD_PAYTO + NUMGUARD_FACILITATOR_URL in the dashboard. With NUMGUARD_PAYTO unset the API runs free (dev mode) so you can test before wiring a wallet. Endpoints: POST /verify_backtest, /verify_model_gap, … ; GET /pricing.

The x402 flow, end to end: the agent POSTs → gets 402 with an accepts block (price, payTo, network) → signs a USDC payment → retries with an X-PAYMENT header → numguard verifies + settles it through the facilitator to your wallet → returns the result. Settlement is the real x402 /verify + /settle handshake (numguard.x402.facilitator_verifier) — facilitator-agnostic: point NUMGUARD_FACILITATOR_URL at any x402 facilitator. Options:

  • Testnet (free, no account): https://x402.org/facilitator with NUMGUARD_NETWORK=base-sepolia — test the whole flow with test-USDC first.
  • Mainnet, self-sovereign: self-host x402-rs (open-source, no third party) and point at your own URL.
  • Mainnet, hosted (non-Coinbase): thirdweb or PayAI facilitators (Base) — set NUMGUARD_FACILITATOR_AUTH if the facilitator needs a key.

2. Serve the MCP server over HTTP (for remote MCP hosts): uvicorn numguard.mcp_server:app (or NUMGUARD_TRANSPORT=streamable-http python -m numguard.mcp_server).

3. Get discovered: server.json (official MCP registry) and smithery.yaml (Smithery) ship in the repo; connect the repo at those registries so agents can find the server. GitHub topics: mcp, mcp-server, x402.

Design notes

  • Statistics are shared with evalgate (zero-dependency); numguard adds the backtest, receipt, metering, and MCP layers on top — it does not re-implement the core checks.
  • Pure-math numerics where possible; cryptography only for Ed25519 receipts (HMAC fallback without it).
  • Every verdict is derived from a computed statistic, never asserted — the same discipline as the book behind it, Measured, Not Believed (leanpub.com/measurednotbelieved).

MIT.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

numguard-0.2.3.tar.gz (133.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

numguard-0.2.3-py3-none-any.whl (127.2 kB view details)

Uploaded Python 3

File details

Details for the file numguard-0.2.3.tar.gz.

File metadata

  • Download URL: numguard-0.2.3.tar.gz
  • Upload date:
  • Size: 133.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for numguard-0.2.3.tar.gz
Algorithm Hash digest
SHA256 b6a283cf868fb0af595257b8ef05e36ccbabc9fe4dda1f5b32f778662f797d7f
MD5 9e0a60654ed355aaf6f55e2225352634
BLAKE2b-256 30cdc826eb9be72cb8999e685b84b7aea4cd033d4d347cd1c3453f7448aa7a9a

See more details on using hashes here.

Provenance

The following attestation bundles were made for numguard-0.2.3.tar.gz:

Publisher: publish.yml on ipezygj/numguard

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file numguard-0.2.3-py3-none-any.whl.

File metadata

  • Download URL: numguard-0.2.3-py3-none-any.whl
  • Upload date:
  • Size: 127.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for numguard-0.2.3-py3-none-any.whl
Algorithm Hash digest
SHA256 b646bcada031cc3a98094f67083d80c8fc71575c5df83f56adbdc11c2b1dfb12
MD5 4fa35bffd1ae6d2ca0084b5cc00a875d
BLAKE2b-256 0bc1dbde01858f394c092e0be512f9a187f0d17d4cd89fda2809454244e0f49b

See more details on using hashes here.

Provenance

The following attestation bundles were made for numguard-0.2.3-py3-none-any.whl:

Publisher: publish.yml on ipezygj/numguard

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.6

2 files

0.2.5

2 files

0.2.4

2 files

This release

0.2.3 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page