Wikidata + Google Knowledge Graph MCP
Token-efficient entity search, resolution and knowledge-graph helpers for AI agents.
Documentation · Install · Demo and benchmark · Changelog
wikidata-google-knowledge-mcp is an open-source MCP server, CLI and Agent Skill for
Claude Code, Cursor, Codex and other MCP clients. It searches and resolves real-world
entities across Wikidata and, optionally, the Google Knowledge Graph without dumping large
provider responses into your LLM context. You get a few ranked candidates, the facts you
asked for, and a deterministic entity-resolution decision with evidence codes, including an
explicit HOLD or AMBIGUOUS when the evidence does not settle the identity.
- Wikidata MCP, without the payload dumps: bounded candidate search (3 by default, 5 at most), selected facts instead of whole items, bounded relationships.
- Entity resolution and entity linking: match your own records (name, kind, city, coordinates, official website, dates, venue, creator...) to Wikidata QIDs and Google KG ids.
- Google Knowledge Graph MCP helpers (optional): exact cross-provider joins
(Google
/m/...= Wikidata P646, Google/g/...= Wikidata P2671). - Batch mode: a streaming, resumable JSONL resolver for thousands of records.
Read-only. Wikidata needs no API key.
Contents
Why this exists · Features · Quick start · MCP installation · Wikidata-only usage · Google Knowledge Graph setup · Tool reference · Entity resolution · Batch resolution · Ambiguity and evidence · Examples · Benchmark · Supported clients · Privacy · Development
Why this exists
Generic Wikidata tools answer "what does Wikidata say about X?". Agents that link records to a knowledge graph mostly need to know which entity is theirs. When a model does that through raw search and statement dumps, it reads dozens of namesakes and full claim lists, then picks one by prose reasoning. This project moves that work into deterministic code and gives the model a short answer.
| Raw Wikidata MCP workflow | wikidata-google-knowledge-mcp | |
|---|---|---|
| Search results | every provider hit goes into context | 3 candidates by default (max 5) with name/place/type match flags |
| Entity facts | full statement lists | only the properties you ask for; ranks, qualifiers and references on request |
| Output size | unbounded | every response is fitted to a byte budget (6,000 bytes by default, configurable) |
| Relationships | hand-written SPARQL | kg_related: outgoing, inverse or class hierarchy, depth- and row-capped |
| Identity decision | the model picks from candidates | deterministic decision plus evidence codes (OFFICIAL_HOST_EXACT, GEO_MATCH...) |
| Namesakes | easy to pick the wrong one | HOLD / AMBIGUOUS are first-class results; no identity probabilities |
| Google Knowledge Graph | separate tool, manual comparison | exact joins via Wikidata P646 / P2671, kept apart from identity proof |
| Thousands of entities | one model tool call per entity | wdkg resolve-batch: streaming JSONL, checkpoint/resume, request budgets |
| Repeated lookups | new provider calls each time | local SQLite cache; a finished batch rerun makes zero provider requests |
Features
- Compact search.
kg_searchreturns at most 5 candidates (3 by default) and says how the name, place and type matched. Place and type hints are checked locally; they are not sent to the provider. - Selected facts.
kg_entityreturns the properties you name (P31,P131...) or a short identity/location overview. Up to three properties can include ranks, qualifiers and references. - Relationship lookup.
kg_relatedfollows one property outward or inward, or walks instance-of/subclass-of, with depth and row caps. - Entity resolution.
kg_resolvetakes a local identity envelope (name, kind, city, country, coordinates, address, official URL, aliases, venue, event date, organizer, role, occupation, creator, year, existing ids) and reconciles provider candidates deterministically. - Cross-provider reconciliation. Google
kg:/m/...ids are joined to Wikidata P646 andkg:/g/...ids to P2671 by exact string equality; canonical Wikipedia URLs are compared with Wikidata sitelinks. Provider agreement is reported separately from proof that the entity is yours. - Ambiguity preservation. Namesakes produce
AMBIGUOUSorHOLD, never a silent pick. Provider ranking is never turned into an identity probability. - Evidence codes. Decisions cite closed-vocabulary codes such as
EXTERNAL_ID_EXACT,OFFICIAL_HOST_EXACT,GEO_MATCH,ADDRESS_MATCH,CITY_MATCH,COUNTRY_MATCH,KIND_COMPATIBLE,EVENT_DATE_MATCHandVENUE_MATCH, not confidence percentages. - Batch resolution.
wdkg resolve-batchstreams JSONL in and out with bounded memory, checkpoint/resume, per-provider request caps, per-row errors, cache reuse and deterministic output, plus an auditable evidence export. - Status and diagnostics.
kg_statusreports provider endpoints, whether a Google key is present and where it came from, cache statistics and limits. It never prints the key.
One Python core serves the MCP server, the wdkg CLI and the Agent Skill; there is no
second resolver implementation.
Quick start
Requirements: uv (it provides Python 3.11+ if needed) and outbound HTTPS.
# run once without installing (from PyPI)
uvx --from wikidata-google-knowledge-mcp==0.2.1 wdkg search "Fox Theatre" --place Atlanta
# or install both commands (wdkg and wikidata-google-knowledge-mcp) on PATH
uv tool install wikidata-google-knowledge-mcp==0.2.1
wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org
The first run downloads the package and its dependencies; later runs use uv's cache. The
same release is also installable straight from the GitLab tag:
uv tool install git+https://gitlab.com/revanalex/wikidata-google-knowledge-mcp@v0.2.1.
wdkg --help lists every command with examples and explains the decisions;
wdkg status shows provider and credential status without any network call.
MCP installation
The MCP server is wikidata-google-knowledge-mcp (stdio). Every client below starts it with
the same command:
uvx wikidata-google-knowledge-mcp@0.2.1
If you installed it with uv tool install, the command is just wikidata-google-knowledge-mcp.
Claude Code
As a plugin (MCP server and skill together):
claude plugin marketplace add https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
claude plugin install wikidata-google-knowledge-mcp@wikidata-google-knowledge-mcp
The same works inside a session with /plugin marketplace add ... and /plugin install ....
MCP server only:
claude mcp add --scope user wikidata-google-knowledge -- \
uvx wikidata-google-knowledge-mcp@0.2.1
Codex
As a plugin (MCP server and skill together):
codex plugin marketplace add https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
codex plugin add wikidata-google-knowledge-mcp@wikidata-google-knowledge-mcp
MCP server only, in ~/.codex/config.toml:
[mcp_servers.wikidata-google-knowledge]
command = "uvx"
args = ["wikidata-google-knowledge-mcp@0.2.1"]
startup_timeout_sec = 60 # the first start downloads the package
env_vars = ["GOOGLE_KNOWLEDGE_GRAPH_API_KEY"] # optional: forward the Google key
or codex mcp add wikidata-google-knowledge -- uvx wikidata-google-knowledge-mcp@0.2.1.
Cursor
Add to ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project):
{
"mcpServers": {
"wikidata-google-knowledge": {
"command": "uvx",
"args": ["wikidata-google-knowledge-mcp@0.2.1"]
}
}
}
To load the full plugin (MCP server and skill) locally, clone this repository into
~/.cursor/plugins/local/wikidata-google-knowledge-mcp and reload the window.
Other MCP clients
Any client that launches stdio servers can use the command above. The repository also ships
portable Agent Plugins manifests (plugin.json, mcp.json,
skills/).
Wikidata-only usage
Everything works without Google. Wikidata lookups use the public Wikidata MCP service on
Wikimedia Cloud (wd-mcp.wmcloud.org) for interactive calls, and the Wikidata Action API and Query Service for batch resolution. No account or
key is needed.
wdkg search "Tate Modern" --place London --type museum
wdkg entity Q193375 --props P31,P131,P17 --evidence P131
wdkg related Q193375 --prop P361
wdkg related Q193375 --hierarchy --depth 2
Optional Google Knowledge Graph setup
Google adds a second, independent candidate source and exact id joins. Without a key it is
never called. With a key, interactive search uses it only when you ask (provider="google",
--provider google, --fallback); the resolver uses it for every row with --providers dual,
and in its default minimal mode only for rows Wikidata left unresolved that carry an
official_url, or when the optional model arbitration is on (the cases where Google can change
the result).
-
Create an API key for the Knowledge Graph Search API (Google prerequisites) and restrict it to that API.
-
Provide it as
GOOGLE_KNOWLEDGE_GRAPH_API_KEYin the environment of the MCP server or CLI, or put one literal line in~/.config/wikidata-google-knowledge-mcp/secrets.env(chmod 600), which also works for GUI clients that do not pass your shell environment:GOOGLE_KNOWLEDGE_GRAPH_API_KEY=your-key-hereThe file is parsed, never executed. The key travels only in the
X-Goog-Api-Keyheader. -
Check with
wdkg statusorkg_status: they reportpresentand the source, never the value.
The Knowledge Graph Search API is an older Google API. Its availability, quotas and terms are set by Google and may change; this project cannot promise it will stay available. All Wikidata features keep working without it.
Tool reference
| MCP tool | Use it for | Key parameters |
|---|---|---|
kg_search |
searching by name | query, place, type, lang, limit (≤5), provider (wikidata/google), fallback |
kg_entity |
reading selected facts of one QID | id, props (≤12 PIDs), evidence (≤3 PIDs with ranks, qualifiers, references), lang |
kg_related |
following relationships | id, prop, inverse, hierarchy, depth (≤3), limit (≤25) |
kg_resolve |
resolving one ambiguous real-world entity | name, kind, city, country, latitude/longitude, address, official_url, aliases, venue, event_date, organizer, role, occupation, affiliation, creator, year, existing_wikidata_qid, existing_google_kg_id, lang, explain |
kg_status |
checking provider and credential status | none; makes no network calls |
All tools are read-only and return one compact JSON object with status, warnings and a
meta block (cache state, upstream request count, bytes).
The wdkg CLI exposes the same core: search, entity, related, sparql (guarded,
read-only, LIMIT enforced), status, cache stats|clear, resolve, resolve-batch,
export-evidence and validate-evidence. Run wdkg <command> --help for flags. Output
fields, warnings, limits and settings are documented in
skills/wikidata-google-knowledge/reference.md.
Entity resolution
kg_resolve (one entity) and wdkg resolve / resolve-batch accept an identity envelope.
Only name is required; give whatever you already know:
{"name": "Fox Theatre", "kind": "place", "city": "Atlanta", "country": "US",
"official_url": "https://www.foxtheatre.org"}
Supported fields: name, kind (person, organization, place, event,
event_series, work, other), city, country, latitude + longitude, address,
official_url, aliases, venue, event_date, organizer, role, occupation,
affiliation, creator, year, existing_wikidata_qid, existing_google_kg_id, lang,
and in the CLI/JSONL only source_id, existing_id_trusted and official_url_reviewed.
The last two assert human review, so the MCP tool does not accept them.
Every result carries a decision:
| decision | meaning |
|---|---|
AUTO_MATCH |
a kind-specific rule passed: a local anchor (official host, coordinates, address, creator, date + venue...) agrees and nothing conflicts |
AMBIGUOUS |
several viable identities remain |
HOLD |
the policy will not decide automatically, or evidence is incomplete |
CONFLICT |
evidence disagrees, for example an existing id is contradicted |
NO_CANDIDATE |
the bounded search found nothing usable (not proof that nothing exists) |
MODEL_MATCH |
optional, off by default: a model picked a supplied candidate (CLI only) |
Kind rules in short: a place needs an official host, coordinates or address match; an
organization needs its official host; a person needs a reviewed official site; a dated event
needs its date plus venue, organizer or coordinates; a series needs its official host or
city plus organizer/venue; a work needs its creator. The full rules, evidence vocabulary and
retention basis are in
skills/wikidata-google-knowledge/resolution.md.
Batch resolution
For corpora, use the CLI rather than repeated MCP calls:
wdkg resolve-batch entities.jsonl --out resolved.jsonl \
--max-wikidata-requests 2000 --max-google-requests 0
- One output record per input line, in input order, plus
resolved.jsonl.receipt.jsonwith decision counts, request usage and stop reason. - The output file is the checkpoint. Rerunning the same command skips finished rows; rows stopped by a budget are redone. A rerun of a finished batch makes zero provider requests.
- Memory is bounded by the chunk size, not the corpus. Provider errors stay per row.
--dry-runparses and plans without network.--providers dualadds Google for every row (needs a key); the defaultminimalmode calls Google only when it can change a result.wdkg export-evidence resolved.jsonl --bundle evidence/writes an auditable bundle (decisions.jsonl,evidence.jsonl,manifest.json,validation.json);wdkg validate-evidence evidence/checks it.
Ambiguity and evidence model
Two kinds of agreement are kept apart:
- Provider concordance: Google and Wikidata describe the same thing
(
EXTERNAL_ID_EXACT,WIKIPEDIA_EXACT,PROVIDER_NAME_AGREEMENT...). This says nothing about whether that thing is your entity. - Local anchors: the external entity agrees with facts you supplied
(
OFFICIAL_HOST_EXACT,GEO_MATCH,ADDRESS_MATCH,CITY_MATCH,EVENT_DATE_MATCH,VENUE_MATCH,CREATOR_MATCH...). Only these can produceAUTO_MATCH.
Conflicts (KIND_CONFLICT, GEO_CONFLICT, EVENT_GRAIN_CONFLICT, CITY_CONFLICT...)
reject a candidate or block automation. Google's resultScore is passed through as
result_score_raw for display and never used in a decision. Nothing is guessed: missing
input stays missing, and QIDs or Google ids only ever come from provider responses or your
input.
Examples
The outputs below come from real runs of v0.1.0 against live Wikidata (and, where noted, the Google Knowledge Graph), shortened to the relevant fields.
Fox Theatre, Atlanta
An exact label lookup finds 131 Wikidata items named "Fox Theatre" (the namesake_heavy:131
warning). With an official website:
wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org --providers dual
{
"decision": "AUTO_MATCH",
"wikidata_qid": "Q1440190",
"google_kg_id": "kg:/m/04qrhq",
"provider_concordance": ["EXTERNAL_ID_EXACT", "PROVIDER_HOST_AGREEMENT", "PROVIDER_NAME_AGREEMENT",
"PROVIDER_TYPE_COMPATIBLE", "WIKIPEDIA_EXACT"],
"local_anchor_evidence": ["CITY_MATCH", "OFFICIAL_HOST_EXACT"],
"evidence_codes": ["CITY_MATCH", "EXTERNAL_ID_EXACT", "KIND_COMPATIBLE", "NAME_EXACT", "OFFICIAL_HOST_EXACT", "..."],
"warnings": ["namesake_heavy:131"]
}
EXTERNAL_ID_EXACT here means Google's kg:/m/04qrhq equals the P646 value on Wikidata
Q1440190. Without Google (--providers minimal, no key) the same call still returns
AUTO_MATCH for Q1440190, with google_kg_id: null.
The same name without an anchor: HOLD
wdkg resolve "Fox Theatre" --kind place --city Atlanta --providers dual
{
"decision": "HOLD",
"reasons": ["NO_LOCAL_ANCHOR"],
"candidate_ids": ["Q1440190", "Q3080199", "Q8565227", "Q559730"],
"candidate_pairs": [
{"wikidata_qid": "Q1440190", "google_kg_id": "kg:/m/04qrhq", "methods": ["EXTERNAL_ID_EXACT", "WIKIPEDIA_EXACT"], "selected": false},
{"wikidata_qid": "Q3080199", "google_kg_id": "kg:/m/05c6w7", "methods": ["EXTERNAL_ID_EXACT", "WIKIPEDIA_EXACT"], "selected": false}
],
"local_anchor_evidence": ["CITY_MATCH"]
}
Google and Wikidata agree exactly on four different Fox Theatres. That agreement is
provider concordance, not proof of which one is yours, and a city name alone is not enough
for a place, so the result is HOLD.
A museum and its building: AMBIGUOUS
wdkg resolve "Tate Modern" --kind place --city London --lat 51.5076 --lon -0.0994
The coordinates match both Tate Modern (Q193375, the gallery) and Bankside Power Station
(Q806832, the building that houses it). The result is AMBIGUOUS
(MULTIPLE_ANCHORED_CANDIDATES) with both ids in candidate_ids, instead of a guess. An
official URL would settle it.
Event occurrence versus event series
wdkg resolve "Primavera Sound" --kind event --date 2019-05-30 --city Barcelona
# decision NO_CANDIDATE, conflict EVENT_GRAIN_CONFLICT on Q2439480 (the festival series)
wdkg resolve "Primavera Sound" --kind event_series --city Barcelona --url https://www.primaverasound.com
# decision AUTO_MATCH, wikidata_qid Q2439480, local anchors CITY_MATCH + OFFICIAL_HOST_EXACT
The 2019 edition is a dated occurrence; Q2439480 is the recurring festival. The resolver never links an occurrence to its series or the other way round.
Batch JSONL
examples/entities.jsonl holds five of the records above plus an
underspecified person:
wdkg resolve-batch examples/entities.jsonl --out resolved.jsonl --max-google-requests 0
{"op": "resolve-batch", "status": "ok", "rows": 6,
"decisions": {"AUTO_MATCH": 2, "AMBIGUOUS": 2, "HOLD": 1, "NO_CANDIDATE": 1, "CONFLICT": 0, "MODEL_MATCH": 0}}
Running the same command again reports "processed": 0, "skipped_already_done": 6 and no
upstream requests.
Benchmark
Measured on seven questions replayed from captured Wikidata MCP responses (offline, CC0 data),
the answer the model reads is 61–85% smaller than the upstream text for searches with many
hits and for entity facts, and larger for tiny answers (a 25-hit search, a one-level class
hierarchy), because match flags, warnings and the meta block cost bytes. A live session
against Wikidata used 3 upstream requests per cold call and 0 on repeat. Method, per-case
numbers, what is omitted and the limitations are on the
benchmark page; there is no accuracy claim.
Supported clients
- Claude Code (plugin or
claude mcp add) - Codex (plugin or
config.toml) - Cursor (
mcp.json, or the plugin loaded locally) - Any MCP client that can launch a stdio server, plus any agent that can run the
wdkgCLI
The skill in skills/wikidata-google-knowledge/
tells the agent which tool to use when; it is plain Markdown and works across clients.
Privacy and provider data
- What leaves your machine: entity names and the parameters of each lookup go to
Wikidata services (the Wikidata MCP service on Wikimedia Cloud, the Wikidata Action API,
the Wikidata Query Service). They go to Google only when you enable it. Place and type
hints for
kg_searchare checked locally and are not sent. There is no telemetry. - What is stored locally: a SQLite cache in
~/.cache/wikidata-google-knowledge-mcp/(Wikidata results expire after 7 days by default;WDKG_CACHE=0or--no-cachedisables it) and, for the resolver, an evidence store in the same directory. - Google content: Google responses are cached only as long as their
Cache-Controlheader allows. The Knowledge Graph Search API currently sends none, so nothing received only from Google is stored: no names, descriptions, scores, website URLs or Wikipedia URLs. A Google id is kept only when it equals an independently obtained string (a Wikidata P646/P2671 value or your own input). - Credentials: the Google key is read from the environment or a local file, sent only in a request header, and redacted from all output.
Details: docs/privacy.md.
Development
git clone https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
cd wikidata-google-knowledge-mcp
uv sync
uv run pytest -q # offline: fixtures and fake transports only
uv run python scripts/check_public_tree.py # forbidden files, secret patterns, manifest checks
uv run --group docs python scripts/build_site.py --out public # documentation site
uv run --group bench python benchmarks/replay_benchmark.py # offline benchmark
The tests never call Wikidata, Google or any model API. CI runs the same checks plus a clean wheel install, a CLI smoke test, an MCP stdio start-up check and a secret scan. See CHANGELOG.md for releases and SECURITY.md for reporting vulnerabilities.
License
MIT; see LICENSE. Wikidata content is available under CC0. Google Knowledge Graph results are subject to Google's terms and are not redistributed by this project.
Non-affiliation
This is an independent open-source project. It is not affiliated with, endorsed by or sponsored by the Wikimedia Foundation, Wikimedia Deutschland, Google, OpenAI, Anthropic or Anysphere (Cursor). Product names are used only to describe compatibility.
Built by the DondeGo.com team as part of our work on semantic city discovery.
Release files for wikidata-google-knowledge-mcp 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| wikidata_google_knowledge_mcp-0.2.1.tar.gz | 288.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| wikidata_google_knowledge_mcp-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 398.8 kB
Release files / wikidata_google_knowledge_mcp-0.2.1.tar.gz
| Download URL | wikidata_google_knowledge_mcp-0.2.1.tar.gz |
|---|---|
| Size | 288.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
88bc2332c27f5c71e023599cf434a44dd25682ce6425713ae57964b041d817d0
|
|
BLAKE2b-256 checksum How to use checksums |
9eba394cd21c6f5d33f29eca23be06a2e0fdb76cf670760a66dbb4fdccbf5085
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitLab CI/CD, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / wikidata_google_knowledge_mcp-0.2.1-py3-none-any.whl
| Download URL | wikidata_google_knowledge_mcp-0.2.1-py3-none-any.whl |
|---|---|
| Size | 110.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
13067de101edc5d39fee248ba962229d312e229172e12a0735e95634e134a368
|
|
BLAKE2b-256 checksum How to use checksums |
f35f85135c8ae3477f2a8bcbccee954517aba3b8b94c4febffdfccc8f8fd20e7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitLab CI/CD, verified by PyPI on Sep 28, 2026.
Transparency log