Skip to main content

tdqs

The Tool Definition Quality Score (TDQS) for Python: a CLI and a library that score how well an MCP tool definition communicates to an AI agent, exactly as the specification defines it. It is the same reference implementation that ships for Node as the tdqs npm package, stage for stage, and the two produce the same numbers, the same hashes and the same prompts.

TDQS scores a definition, not behaviour. The inputs are what an MCP client sees from tools/list — name, title, description, input schema, output schema, annotations — and the output is a score from 1.0 to 5.0 with a letter tier, per tool and per server, with a justification for every dimension.

Install

pip install tdqs
# or run it without installing
uvx tdqs --help

Python 3.10 or newer.

Lint: deterministic, no model, no key

tdqs lint --file tools.json
tdqs lint --command "uvx my-mcp-server"
tdqs lint --url https://mcp.example.com/mcp --header "Authorization: Bearer …"

lint runs the stages of the pipeline that need no model: the context signals (parameter counts, schema description coverage, annotation values, invocation cost, the definition's hash and byte size), the hard gates (no description, tautological description), the shadow prefilter across the tool set, and the checklist the specification ranks highest. It exits 1 on an error-level finding, which makes it a pull request check:

tdqs lint --file tools.json --fail-on warning --format markdown --output tdqs-lint.md

A lint finding names a fix. It is not a score, and it never pretends to be one.

Score: the full rubric

export TDQS_BASE_URL=https://api.openai.com/v1   # any OpenAI-compatible endpoint
export TDQS_API_KEY=…
export TDQS_MODEL=…

tdqs score --file tools.json
tdqs score --command "uvx my-mcp-server" --fail-under B --format markdown

score sends every tool through the rubric (six dimensions, 1–5 each, with the specification's system prompt verbatim), runs the server coherence evaluation (four dimensions plus shadowing-risk confirmation), and rolls both up into the server score with integer arithmetic. The report is stamped with the specification version and the model, because a score is calibrated to a rubric+model pair and is not comparable to anything without both.

Turn extended reasoning off. The reference model reasons before it answers unless told not to, which makes a call take a minute instead of seconds — and the specification's calibration examples reproduce with reasoning off. How to say so is provider-specific, so it is an opaque JSON object merged into every request:

tdqs score --file tools.json --request-overrides '{"reasoning":{"enabled":false}}'   # OpenRouter
# or TDQS_REQUEST_OVERRIDES in the environment; DeepSeek directly takes {"thinking":{"type":"disabled"}}

--hosted https://tdqs.dev scores through a hosted TDQS site instead of a model key of your own, and prints the report's URL. It takes that site's API key as --api-key or TDQS_API_KEY; the site's account page is where keys come from.

Input is exactly one of --file (a tools/list result, an array of tools, or a single tool; - reads stdin), --command (a stdio server) or --url (a Streamable HTTP server).

Exit code Meaning
0 done
1 the threshold was not met (--fail-on for lint, --fail-under for score)
2 usage error, unreadable input, unreachable server, or a model failure

--format is text (default), markdown or json. The JSON formats are the ones the npm package publishes as JSON Schema, schemas/score-report.json and schemas/lint-report.json.

Library

from tdqs import create_llm_client, lint_server, parse_tool_definitions, score_server

parsed = parse_tool_definitions(response.json())
server_name = parsed.server_name or "my-server"

# No model involved.
lint = lint_server(server_name=server_name, tools=parsed.tools)

# The full pipeline.
report = score_server(
    llm=create_llm_client(
        api_key=api_key,
        base_url=base_url,
        model=model,
        request_overrides={"reasoning": {"enabled": False}},
    ),
    server_name=server_name,
    tools=parsed.tools,
)

report.server_score.overall_tier  # "A" | "B" | "C" | "D" | "F"
report.tools[0].justifications["usage_guidelines"]  # Justification(score=…, justification=…)
report.model_dump()  # the specification's JSON, camelCase keys

Every stage is exported on its own — compute_context_signals, evaluate_hard_gates, compute_tdqs, find_shadow_candidates, build_tool_scoring_prompt, score_tool_definition, score_server_coherence, rollup_server_score — along with the specification's metadata (TOOL_DIMENSIONS, COHERENCE_DIMENSIONS, FLAGS, TIERS, LINT_RULES, SPEC_VERSION) and the two system prompts, so a registry or a gateway can build on the same pieces. Reports are pydantic models; model_dump() is the specification's JSON and ScoreReport.model_validate() reads it back.

request_hosted_report(...) is the hosted mode as a function: it submits the definitions to a TDQS site and polls until the report is done.

What is deterministic and what is not

Stages 1, 2 and 4 of the pipeline, the shadow prefilter, and every rollup are deterministic and reproducible from the definitions alone; inputHash is computed the same way the Glama registry computes it, so a hash here matches the one on a server's public score page. Stage 3 — the rubric — and the coherence evaluation are model calls. The specification pins the prompts, the output contract and the calibration examples; the model is the remaining variable, which is why every report names it. Swap models and expect to re-score.

Specification

This package follows TDQS 1.2. The prompts are compared byte for byte against the specification in the test suite, and the deterministic stages are compared against the Node reference implementation's fixtures, so the implementations cannot drift apart silently.

Metadata

Release files for tdqs 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tdqs 0.1.0
File Size Uploaded
tdqs-0.1.0.tar.gz 50.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tdqs 0.1.0
File Interpreter ABI Platform
tdqs-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 122.3 kB

Release files / tdqs-0.1.0.tar.gz

Download URL tdqs-0.1.0.tar.gz
Size 50.3 kB
Tags Source
SHA-256 checksum
How to use checksums
d20b6bf8c49c4e05b43b5ed2fbd5d2ae1fd2eee2dbc3b61dada57d31c8262f15
BLAKE2b-256 checksum
How to use checksums
02ce9028524d8a59626f705220431447cd5c221e1305015e9d648176f2f1f001
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release files / tdqs-0.1.0-py3-none-any.whl

Download URL tdqs-0.1.0-py3-none-any.whl
Size 72.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7acdd5dc2edfe892b647d1f9f64caf672c2c4e9836ed2849ebbd59725dc91f11
BLAKE2b-256 checksum
How to use checksums
2364abc0726a3f7da232125ea5add3e2ee19e1bb7c8d23143e80d7ff43fe4f60
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page