Skip to main content

Wikidata + Google Knowledge Graph MCP

Token-efficient entity search, resolution and knowledge-graph helpers for AI agents.

Documentation · Install · Demo and benchmark · Changelog

wikidata-google-knowledge-mcp is an open-source MCP server, CLI and Agent Skill for Claude Code, Cursor, Codex and other MCP clients. It searches and resolves real-world entities across Wikidata and, optionally, the Google Knowledge Graph without dumping large provider responses into your LLM context. You get a few ranked candidates, the facts you asked for, and a deterministic entity-resolution decision with evidence codes, including an explicit HOLD or AMBIGUOUS when the evidence does not settle the identity.

  • Wikidata MCP, without the payload dumps: bounded candidate search (3 by default, 5 at most), selected facts instead of whole items, bounded relationships.
  • Entity resolution and entity linking: match your own records (name, kind, city, coordinates, official website, dates, venue, creator...) to Wikidata QIDs and Google KG ids.
  • Google Knowledge Graph MCP helpers (optional): exact cross-provider joins (Google /m/... = Wikidata P646, Google /g/... = Wikidata P2671).
  • Batch mode: a streaming, resumable JSONL resolver for thousands of records.

Read-only. Wikidata needs no API key.

Contents

Why this exists · Features · Quick start · MCP installation · Wikidata-only usage · Google Knowledge Graph setup · Tool reference · Entity resolution · Batch resolution · Ambiguity and evidence · Examples · Benchmark · Supported clients · Privacy · Development

Why this exists

Generic Wikidata tools answer "what does Wikidata say about X?". Agents that link records to a knowledge graph mostly need to know which entity is theirs. When a model does that through raw search and statement dumps, it reads dozens of namesakes and full claim lists, then picks one by prose reasoning. This project moves that work into deterministic code and gives the model a short answer.

Raw Wikidata MCP workflow wikidata-google-knowledge-mcp
Search results every provider hit goes into context 3 candidates by default (max 5) with name/place/type match flags
Entity facts full statement lists only the properties you ask for; ranks, qualifiers and references on request
Output size unbounded every response is fitted to a byte budget (6,000 bytes by default, configurable)
Relationships hand-written SPARQL kg_related: outgoing, inverse or class hierarchy, depth- and row-capped
Identity decision the model picks from candidates deterministic decision plus evidence codes (OFFICIAL_HOST_EXACT, GEO_MATCH...)
Namesakes easy to pick the wrong one HOLD / AMBIGUOUS are first-class results; no identity probabilities
Google Knowledge Graph separate tool, manual comparison exact joins via Wikidata P646 / P2671, kept apart from identity proof
Thousands of entities one model tool call per entity wdkg resolve-batch: streaming JSONL, checkpoint/resume, request budgets
Repeated lookups new provider calls each time local SQLite cache; a finished batch rerun makes zero provider requests

Features

  1. Compact search. kg_search returns at most 5 candidates (3 by default) and says how the name, place and type matched. Place and type hints are checked locally; they are not sent to the provider.
  2. Selected facts. kg_entity returns the properties you name (P31, P131...) or a short identity/location overview. Up to three properties can include ranks, qualifiers and references.
  3. Relationship lookup. kg_related follows one property outward or inward, or walks instance-of/subclass-of, with depth and row caps.
  4. Entity resolution. kg_resolve takes a local identity envelope (name, kind, city, country, coordinates, address, official URL, aliases, venue, event date, organizer, role, occupation, creator, year, existing ids) and reconciles provider candidates deterministically.
  5. Cross-provider reconciliation. Google kg:/m/... ids are joined to Wikidata P646 and kg:/g/... ids to P2671 by exact string equality; canonical Wikipedia URLs are compared with Wikidata sitelinks. Provider agreement is reported separately from proof that the entity is yours.
  6. Ambiguity preservation. Namesakes produce AMBIGUOUS or HOLD, never a silent pick. Provider ranking is never turned into an identity probability.
  7. Evidence codes. Decisions cite closed-vocabulary codes such as EXTERNAL_ID_EXACT, OFFICIAL_HOST_EXACT, GEO_MATCH, ADDRESS_MATCH, CITY_MATCH, COUNTRY_MATCH, KIND_COMPATIBLE, EVENT_DATE_MATCH and VENUE_MATCH, not confidence percentages.
  8. Batch resolution. wdkg resolve-batch streams JSONL in and out with bounded memory, checkpoint/resume, per-provider request caps, per-row errors, cache reuse and deterministic output, plus an auditable evidence export.
  9. Status and diagnostics. kg_status reports provider endpoints, whether a Google key is present and where it came from, cache statistics and limits. It never prints the key.

One Python core serves the MCP server, the wdkg CLI and the Agent Skill; there is no second resolver implementation.

Quick start

Requirements: uv (it provides Python 3.11+ if needed) and outbound HTTPS.

# run once without installing (from PyPI)
uvx --from wikidata-google-knowledge-mcp==0.2.1 wdkg search "Fox Theatre" --place Atlanta

# or install both commands (wdkg and wikidata-google-knowledge-mcp) on PATH
uv tool install wikidata-google-knowledge-mcp==0.2.1
wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org

The first run downloads the package and its dependencies; later runs use uv's cache. The same release is also installable straight from the GitLab tag: uv tool install git+https://gitlab.com/revanalex/wikidata-google-knowledge-mcp@v0.2.1. wdkg --help lists every command with examples and explains the decisions; wdkg status shows provider and credential status without any network call.

MCP installation

The MCP server is wikidata-google-knowledge-mcp (stdio). Every client below starts it with the same command:

uvx wikidata-google-knowledge-mcp@0.2.1

If you installed it with uv tool install, the command is just wikidata-google-knowledge-mcp.

Claude Code

As a plugin (MCP server and skill together):

claude plugin marketplace add https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
claude plugin install wikidata-google-knowledge-mcp@wikidata-google-knowledge-mcp

The same works inside a session with /plugin marketplace add ... and /plugin install .... MCP server only:

claude mcp add --scope user wikidata-google-knowledge -- \
  uvx wikidata-google-knowledge-mcp@0.2.1

Codex

As a plugin (MCP server and skill together):

codex plugin marketplace add https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
codex plugin add wikidata-google-knowledge-mcp@wikidata-google-knowledge-mcp

MCP server only, in ~/.codex/config.toml:

[mcp_servers.wikidata-google-knowledge]
command = "uvx"
args = ["wikidata-google-knowledge-mcp@0.2.1"]
startup_timeout_sec = 60                     # the first start downloads the package
env_vars = ["GOOGLE_KNOWLEDGE_GRAPH_API_KEY"]  # optional: forward the Google key

or codex mcp add wikidata-google-knowledge -- uvx wikidata-google-knowledge-mcp@0.2.1.

Cursor

Add to ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project):

{
  "mcpServers": {
    "wikidata-google-knowledge": {
      "command": "uvx",
      "args": ["wikidata-google-knowledge-mcp@0.2.1"]
    }
  }
}

To load the full plugin (MCP server and skill) locally, clone this repository into ~/.cursor/plugins/local/wikidata-google-knowledge-mcp and reload the window.

Other MCP clients

Any client that launches stdio servers can use the command above. The repository also ships portable Agent Plugins manifests (plugin.json, mcp.json, skills/).

Wikidata-only usage

Everything works without Google. Wikidata lookups use the public Wikidata MCP service on Wikimedia Cloud (wd-mcp.wmcloud.org) for interactive calls, and the Wikidata Action API and Query Service for batch resolution. No account or key is needed.

wdkg search "Tate Modern" --place London --type museum
wdkg entity Q193375 --props P31,P131,P17 --evidence P131
wdkg related Q193375 --prop P361
wdkg related Q193375 --hierarchy --depth 2

Optional Google Knowledge Graph setup

Google adds a second, independent candidate source and exact id joins. Without a key it is never called. With a key, interactive search uses it only when you ask (provider="google", --provider google, --fallback); the resolver uses it for every row with --providers dual, and in its default minimal mode only for rows Wikidata left unresolved that carry an official_url, or when the optional model arbitration is on (the cases where Google can change the result).

  1. Create an API key for the Knowledge Graph Search API (Google prerequisites) and restrict it to that API.

  2. Provide it as GOOGLE_KNOWLEDGE_GRAPH_API_KEY in the environment of the MCP server or CLI, or put one literal line in ~/.config/wikidata-google-knowledge-mcp/secrets.env (chmod 600), which also works for GUI clients that do not pass your shell environment:

    GOOGLE_KNOWLEDGE_GRAPH_API_KEY=your-key-here
    

    The file is parsed, never executed. The key travels only in the X-Goog-Api-Key header.

  3. Check with wdkg status or kg_status: they report present and the source, never the value.

The Knowledge Graph Search API is an older Google API. Its availability, quotas and terms are set by Google and may change; this project cannot promise it will stay available. All Wikidata features keep working without it.

Tool reference

MCP tool Use it for Key parameters
kg_search searching by name query, place, type, lang, limit (≤5), provider (wikidata/google), fallback
kg_entity reading selected facts of one QID id, props (≤12 PIDs), evidence (≤3 PIDs with ranks, qualifiers, references), lang
kg_related following relationships id, prop, inverse, hierarchy, depth (≤3), limit (≤25)
kg_resolve resolving one ambiguous real-world entity name, kind, city, country, latitude/longitude, address, official_url, aliases, venue, event_date, organizer, role, occupation, affiliation, creator, year, existing_wikidata_qid, existing_google_kg_id, lang, explain
kg_status checking provider and credential status none; makes no network calls

All tools are read-only and return one compact JSON object with status, warnings and a meta block (cache state, upstream request count, bytes).

The wdkg CLI exposes the same core: search, entity, related, sparql (guarded, read-only, LIMIT enforced), status, cache stats|clear, resolve, resolve-batch, export-evidence and validate-evidence. Run wdkg <command> --help for flags. Output fields, warnings, limits and settings are documented in skills/wikidata-google-knowledge/reference.md.

Entity resolution

kg_resolve (one entity) and wdkg resolve / resolve-batch accept an identity envelope. Only name is required; give whatever you already know:

{"name": "Fox Theatre", "kind": "place", "city": "Atlanta", "country": "US",
 "official_url": "https://www.foxtheatre.org"}

Supported fields: name, kind (person, organization, place, event, event_series, work, other), city, country, latitude + longitude, address, official_url, aliases, venue, event_date, organizer, role, occupation, affiliation, creator, year, existing_wikidata_qid, existing_google_kg_id, lang, and in the CLI/JSONL only source_id, existing_id_trusted and official_url_reviewed. The last two assert human review, so the MCP tool does not accept them.

Every result carries a decision:

decision meaning
AUTO_MATCH a kind-specific rule passed: a local anchor (official host, coordinates, address, creator, date + venue...) agrees and nothing conflicts
AMBIGUOUS several viable identities remain
HOLD the policy will not decide automatically, or evidence is incomplete
CONFLICT evidence disagrees, for example an existing id is contradicted
NO_CANDIDATE the bounded search found nothing usable (not proof that nothing exists)
MODEL_MATCH optional, off by default: a model picked a supplied candidate (CLI only)

Kind rules in short: a place needs an official host, coordinates or address match; an organization needs its official host; a person needs a reviewed official site; a dated event needs its date plus venue, organizer or coordinates; a series needs its official host or city plus organizer/venue; a work needs its creator. The full rules, evidence vocabulary and retention basis are in skills/wikidata-google-knowledge/resolution.md.

Batch resolution

For corpora, use the CLI rather than repeated MCP calls:

wdkg resolve-batch entities.jsonl --out resolved.jsonl \
    --max-wikidata-requests 2000 --max-google-requests 0
  • One output record per input line, in input order, plus resolved.jsonl.receipt.json with decision counts, request usage and stop reason.
  • The output file is the checkpoint. Rerunning the same command skips finished rows; rows stopped by a budget are redone. A rerun of a finished batch makes zero provider requests.
  • Memory is bounded by the chunk size, not the corpus. Provider errors stay per row.
  • --dry-run parses and plans without network. --providers dual adds Google for every row (needs a key); the default minimal mode calls Google only when it can change a result.
  • wdkg export-evidence resolved.jsonl --bundle evidence/ writes an auditable bundle (decisions.jsonl, evidence.jsonl, manifest.json, validation.json); wdkg validate-evidence evidence/ checks it.

Ambiguity and evidence model

Two kinds of agreement are kept apart:

  • Provider concordance: Google and Wikidata describe the same thing (EXTERNAL_ID_EXACT, WIKIPEDIA_EXACT, PROVIDER_NAME_AGREEMENT...). This says nothing about whether that thing is your entity.
  • Local anchors: the external entity agrees with facts you supplied (OFFICIAL_HOST_EXACT, GEO_MATCH, ADDRESS_MATCH, CITY_MATCH, EVENT_DATE_MATCH, VENUE_MATCH, CREATOR_MATCH...). Only these can produce AUTO_MATCH.

Conflicts (KIND_CONFLICT, GEO_CONFLICT, EVENT_GRAIN_CONFLICT, CITY_CONFLICT...) reject a candidate or block automation. Google's resultScore is passed through as result_score_raw for display and never used in a decision. Nothing is guessed: missing input stays missing, and QIDs or Google ids only ever come from provider responses or your input.

Examples

The outputs below come from real runs of v0.1.0 against live Wikidata (and, where noted, the Google Knowledge Graph), shortened to the relevant fields.

Fox Theatre, Atlanta

An exact label lookup finds 131 Wikidata items named "Fox Theatre" (the namesake_heavy:131 warning). With an official website:

wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org --providers dual
{
  "decision": "AUTO_MATCH",
  "wikidata_qid": "Q1440190",
  "google_kg_id": "kg:/m/04qrhq",
  "provider_concordance": ["EXTERNAL_ID_EXACT", "PROVIDER_HOST_AGREEMENT", "PROVIDER_NAME_AGREEMENT",
                           "PROVIDER_TYPE_COMPATIBLE", "WIKIPEDIA_EXACT"],
  "local_anchor_evidence": ["CITY_MATCH", "OFFICIAL_HOST_EXACT"],
  "evidence_codes": ["CITY_MATCH", "EXTERNAL_ID_EXACT", "KIND_COMPATIBLE", "NAME_EXACT", "OFFICIAL_HOST_EXACT", "..."],
  "warnings": ["namesake_heavy:131"]
}

EXTERNAL_ID_EXACT here means Google's kg:/m/04qrhq equals the P646 value on Wikidata Q1440190. Without Google (--providers minimal, no key) the same call still returns AUTO_MATCH for Q1440190, with google_kg_id: null.

The same name without an anchor: HOLD

wdkg resolve "Fox Theatre" --kind place --city Atlanta --providers dual
{
  "decision": "HOLD",
  "reasons": ["NO_LOCAL_ANCHOR"],
  "candidate_ids": ["Q1440190", "Q3080199", "Q8565227", "Q559730"],
  "candidate_pairs": [
    {"wikidata_qid": "Q1440190", "google_kg_id": "kg:/m/04qrhq", "methods": ["EXTERNAL_ID_EXACT", "WIKIPEDIA_EXACT"], "selected": false},
    {"wikidata_qid": "Q3080199", "google_kg_id": "kg:/m/05c6w7", "methods": ["EXTERNAL_ID_EXACT", "WIKIPEDIA_EXACT"], "selected": false}
  ],
  "local_anchor_evidence": ["CITY_MATCH"]
}

Google and Wikidata agree exactly on four different Fox Theatres. That agreement is provider concordance, not proof of which one is yours, and a city name alone is not enough for a place, so the result is HOLD.

A museum and its building: AMBIGUOUS

wdkg resolve "Tate Modern" --kind place --city London --lat 51.5076 --lon -0.0994

The coordinates match both Tate Modern (Q193375, the gallery) and Bankside Power Station (Q806832, the building that houses it). The result is AMBIGUOUS (MULTIPLE_ANCHORED_CANDIDATES) with both ids in candidate_ids, instead of a guess. An official URL would settle it.

Event occurrence versus event series

wdkg resolve "Primavera Sound" --kind event --date 2019-05-30 --city Barcelona
# decision NO_CANDIDATE, conflict EVENT_GRAIN_CONFLICT on Q2439480 (the festival series)

wdkg resolve "Primavera Sound" --kind event_series --city Barcelona --url https://www.primaverasound.com
# decision AUTO_MATCH, wikidata_qid Q2439480, local anchors CITY_MATCH + OFFICIAL_HOST_EXACT

The 2019 edition is a dated occurrence; Q2439480 is the recurring festival. The resolver never links an occurrence to its series or the other way round.

Batch JSONL

examples/entities.jsonl holds five of the records above plus an underspecified person:

wdkg resolve-batch examples/entities.jsonl --out resolved.jsonl --max-google-requests 0
{"op": "resolve-batch", "status": "ok", "rows": 6,
 "decisions": {"AUTO_MATCH": 2, "AMBIGUOUS": 2, "HOLD": 1, "NO_CANDIDATE": 1, "CONFLICT": 0, "MODEL_MATCH": 0}}

Running the same command again reports "processed": 0, "skipped_already_done": 6 and no upstream requests.

Benchmark

Measured on seven questions replayed from captured Wikidata MCP responses (offline, CC0 data), the answer the model reads is 61–85% smaller than the upstream text for searches with many hits and for entity facts, and larger for tiny answers (a 25-hit search, a one-level class hierarchy), because match flags, warnings and the meta block cost bytes. A live session against Wikidata used 3 upstream requests per cold call and 0 on repeat. Method, per-case numbers, what is omitted and the limitations are on the benchmark page; there is no accuracy claim.

Supported clients

  • Claude Code (plugin or claude mcp add)
  • Codex (plugin or config.toml)
  • Cursor (mcp.json, or the plugin loaded locally)
  • Any MCP client that can launch a stdio server, plus any agent that can run the wdkg CLI

The skill in skills/wikidata-google-knowledge/ tells the agent which tool to use when; it is plain Markdown and works across clients.

Privacy and provider data

  • What leaves your machine: entity names and the parameters of each lookup go to Wikidata services (the Wikidata MCP service on Wikimedia Cloud, the Wikidata Action API, the Wikidata Query Service). They go to Google only when you enable it. Place and type hints for kg_search are checked locally and are not sent. There is no telemetry.
  • What is stored locally: a SQLite cache in ~/.cache/wikidata-google-knowledge-mcp/ (Wikidata results expire after 7 days by default; WDKG_CACHE=0 or --no-cache disables it) and, for the resolver, an evidence store in the same directory.
  • Google content: Google responses are cached only as long as their Cache-Control header allows. The Knowledge Graph Search API currently sends none, so nothing received only from Google is stored: no names, descriptions, scores, website URLs or Wikipedia URLs. A Google id is kept only when it equals an independently obtained string (a Wikidata P646/P2671 value or your own input).
  • Credentials: the Google key is read from the environment or a local file, sent only in a request header, and redacted from all output.

Details: docs/privacy.md.

Development

git clone https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
cd wikidata-google-knowledge-mcp
uv sync
uv run pytest -q                          # offline: fixtures and fake transports only
uv run python scripts/check_public_tree.py  # forbidden files, secret patterns, manifest checks
uv run --group docs python scripts/build_site.py --out public   # documentation site
uv run --group bench python benchmarks/replay_benchmark.py      # offline benchmark

The tests never call Wikidata, Google or any model API. CI runs the same checks plus a clean wheel install, a CLI smoke test, an MCP stdio start-up check and a secret scan. See CHANGELOG.md for releases and SECURITY.md for reporting vulnerabilities.

License

MIT; see LICENSE. Wikidata content is available under CC0. Google Knowledge Graph results are subject to Google's terms and are not redistributed by this project.

Non-affiliation

This is an independent open-source project. It is not affiliated with, endorsed by or sponsored by the Wikimedia Foundation, Wikimedia Deutschland, Google, OpenAI, Anthropic or Anysphere (Cursor). Product names are used only to describe compatibility.


Built by the DondeGo.com team as part of our work on semantic city discovery.

Release files for wikidata-google-knowledge-mcp 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for wikidata-google-knowledge-mcp 0.2.1
File Size Uploaded
wikidata_google_knowledge_mcp-0.2.1.tar.gz 288.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for wikidata-google-knowledge-mcp 0.2.1
File Interpreter ABI Platform
wikidata_google_knowledge_mcp-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 398.8 kB

Release files / wikidata_google_knowledge_mcp-0.2.1.tar.gz

Download URL wikidata_google_knowledge_mcp-0.2.1.tar.gz
Size 288.0 kB
Tags Source
SHA-256 checksum
How to use checksums
88bc2332c27f5c71e023599cf434a44dd25682ce6425713ae57964b041d817d0
BLAKE2b-256 checksum
How to use checksums
9eba394cd21c6f5d33f29eca23be06a2e0fdb76cf670760a66dbb4fdccbf5085
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitLab CI/CD, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / wikidata_google_knowledge_mcp-0.2.1-py3-none-any.whl

Download URL wikidata_google_knowledge_mcp-0.2.1-py3-none-any.whl
Size 110.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
13067de101edc5d39fee248ba962229d312e229172e12a0735e95634e134a368
BLAKE2b-256 checksum
How to use checksums
f35f85135c8ae3477f2a8bcbccee954517aba3b8b94c4febffdfccc8f8fd20e7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitLab CI/CD, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page