tdqs
The Tool Definition Quality Score (TDQS) for Python: a CLI and a library that score how well an MCP tool definition communicates to an AI agent, exactly as the specification defines it. It is the same reference implementation that ships for Node as the tdqs npm package, stage for stage, and the two produce the same numbers, the same hashes and the same prompts.
TDQS scores a definition, not behaviour. The inputs are what an MCP client sees from tools/list — name, title, description, input schema, output schema, annotations — and the output is a score from 1.0 to 5.0 with a letter tier, per tool and per server, with a justification for every dimension.
Install
pip install tdqs
# or run it without installing
uvx tdqs --help
Python 3.10 or newer.
Lint: deterministic, no model, no key
tdqs lint --file tools.json
tdqs lint --command "uvx my-mcp-server"
tdqs lint --url https://mcp.example.com/mcp --header "Authorization: Bearer …"
lint runs the stages of the pipeline that need no model: the context signals (parameter counts, schema description coverage, annotation values, invocation cost, the definition's hash and byte size), the hard gates (no description, tautological description), the shadow prefilter across the tool set, and the checklist the specification ranks highest. It exits 1 on an error-level finding, which makes it a pull request check:
tdqs lint --file tools.json --fail-on warning --format markdown --output tdqs-lint.md
A lint finding names a fix. It is not a score, and it never pretends to be one.
Score: the full rubric
export TDQS_BASE_URL=https://api.openai.com/v1 # any OpenAI-compatible endpoint
export TDQS_API_KEY=…
export TDQS_MODEL=…
tdqs score --file tools.json
tdqs score --command "uvx my-mcp-server" --fail-under B --format markdown
score sends every tool through the rubric (six dimensions, 1–5 each, with the specification's system prompt verbatim), runs the server coherence evaluation (four dimensions plus shadowing-risk confirmation), and rolls both up into the server score with integer arithmetic. The report is stamped with the specification version and the model, because a score is calibrated to a rubric+model pair and is not comparable to anything without both.
Turn extended reasoning off. The reference model reasons before it answers unless told not to, which makes a call take a minute instead of seconds — and the specification's calibration examples reproduce with reasoning off. How to say so is provider-specific, so it is an opaque JSON object merged into every request:
tdqs score --file tools.json --request-overrides '{"reasoning":{"enabled":false}}' # OpenRouter
# or TDQS_REQUEST_OVERRIDES in the environment; DeepSeek directly takes {"thinking":{"type":"disabled"}}
--hosted https://tdqs.dev scores through a hosted TDQS site instead of a model key of your own, and prints the report's URL. It takes that site's API key as --api-key or TDQS_API_KEY; the site's account page is where keys come from.
Input is exactly one of --file (a tools/list result, an array of tools, or a single tool; - reads stdin), --command (a stdio server) or --url (a Streamable HTTP server).
| Exit code | Meaning |
|---|---|
0 |
done |
1 |
the threshold was not met (--fail-on for lint, --fail-under for score) |
2 |
usage error, unreadable input, unreachable server, or a model failure |
--format is text (default), markdown or json. The JSON formats are the ones the npm package publishes as JSON Schema, schemas/score-report.json and schemas/lint-report.json.
Library
from tdqs import create_llm_client, lint_server, parse_tool_definitions, score_server
parsed = parse_tool_definitions(response.json())
server_name = parsed.server_name or "my-server"
# No model involved.
lint = lint_server(server_name=server_name, tools=parsed.tools)
# The full pipeline.
report = score_server(
llm=create_llm_client(
api_key=api_key,
base_url=base_url,
model=model,
request_overrides={"reasoning": {"enabled": False}},
),
server_name=server_name,
tools=parsed.tools,
)
report.server_score.overall_tier # "A" | "B" | "C" | "D" | "F"
report.tools[0].justifications["usage_guidelines"] # Justification(score=…, justification=…)
report.model_dump() # the specification's JSON, camelCase keys
Every stage is exported on its own — compute_context_signals, evaluate_hard_gates, compute_tdqs, find_shadow_candidates, build_tool_scoring_prompt, score_tool_definition, score_server_coherence, rollup_server_score — along with the specification's metadata (TOOL_DIMENSIONS, COHERENCE_DIMENSIONS, FLAGS, TIERS, LINT_RULES, SPEC_VERSION) and the two system prompts, so a registry or a gateway can build on the same pieces. Reports are pydantic models; model_dump() is the specification's JSON and ScoreReport.model_validate() reads it back.
request_hosted_report(...) is the hosted mode as a function: it submits the definitions to a TDQS site and polls until the report is done.
What is deterministic and what is not
Stages 1, 2 and 4 of the pipeline, the shadow prefilter, and every rollup are deterministic and reproducible from the definitions alone; inputHash is computed the same way the Glama registry computes it, so a hash here matches the one on a server's public score page. Stage 3 — the rubric — and the coherence evaluation are model calls. The specification pins the prompts, the output contract and the calibration examples; the model is the remaining variable, which is why every report names it. Swap models and expect to re-score.
Specification
This package follows TDQS 1.2. The prompts are compared byte for byte against the specification in the test suite, and the deterministic stages are compared against the Node reference implementation's fixtures, so the implementations cannot drift apart silently.
Metadata
Release files for tdqs 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tdqs-0.1.0.tar.gz | 50.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tdqs-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 122.3 kB
Release files / tdqs-0.1.0.tar.gz
| Download URL | tdqs-0.1.0.tar.gz |
|---|---|
| Size | 50.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d20b6bf8c49c4e05b43b5ed2fbd5d2ae1fd2eee2dbc3b61dada57d31c8262f15
|
|
BLAKE2b-256 checksum How to use checksums |
02ce9028524d8a59626f705220431447cd5c221e1305015e9d648176f2f1f001
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.
Transparency logRelease files / tdqs-0.1.0-py3-none-any.whl
| Download URL | tdqs-0.1.0-py3-none-any.whl |
|---|---|
| Size | 72.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7acdd5dc2edfe892b647d1f9f64caf672c2c4e9836ed2849ebbd59725dc91f11
|
|
BLAKE2b-256 checksum How to use checksums |
2364abc0726a3f7da232125ea5add3e2ee19e1bb7c8d23143e80d7ff43fe4f60
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.
Transparency log