Wikidata NER Classifier 0.8.0
wikidata-ner-classifier predicts retrieval-oriented NER types in two ways:
- Wikidata items: deterministic prediction from P31/P279 token clues and an optional description.
- Input data: LLM prediction from a target mention and its context, supplied as free text or tabular data.
Both paths return types from the same hierarchy:
coarse_type -> fine_type -> subtype -> specific_type
The prediction can be used to narrow the candidate-retrieval space before entity linking. The library does not make the final identity decision. LLM predictions may include unverified Wikipedia and DBpedia URLs as search hints; a downstream linker must retrieve and verify them. Wikidata QIDs and URLs are deliberately excluded because they would be unique-identity guesses.
Installation
pip install wikidata-ner-classifier
1. Predict NER types for Wikidata items
Use WikidataNERClassifier when the input is already a Wikidata item and its
P31/P279 type labels or aliases are available. Prediction is deterministic and
does not require an LLM or network request.
from wikidata_ner import WikidataNERClassifier
classifier = WikidataNERClassifier()
prediction = classifier.predict(
qid="Q3441181",
types=[
{"id": "Q11424", "name": "film"},
],
description="1964 sword-and-sandal film directed by Giuseppe Vari",
)
print(prediction.coarse_type) # CREATIVE_WORK
print(prediction.fine_type) # FILM
print(prediction.specific_type) # SWORD_AND_SANDAL_FILM
print(prediction.retrieval_key)
# CREATIVE_WORK/FILM/SWORD_AND_SANDAL_FILM
The P31/P279 token clues select the semantic branch. The description may refine the result, but it does not replace the type-token evidence.
A complete entity mapping can also be passed directly:
prediction = classifier.predict_entity(
{
"qid": "Q3441181",
"types": [{"id": "Q11424", "name": "film"}],
"description": "1964 sword-and-sandal film",
}
)
For multiple Wikidata items, use predict_batch():
predictions = classifier.predict_batch(
[
{
"qid": "Q3441181",
"types": [{"name": "film"}],
"description": "1964 sword-and-sandal film",
},
{
"qid": "Q7259",
"types": [{"name": "human"}],
"description": "English mathematician and writer",
},
]
)
2. Predict NER types for input data with an LLM
Use OpenRouterNERClassifier when the input is a mention whose type must be
inferred from context. The context can be free text, a structured record, or a
table cell.
Set an OpenRouter API key:
export OPENROUTER_API_KEY="..."
Create the classifier:
from wikidata_ner import OpenRouterNERClassifier
classifier = OpenRouterNERClassifier(
model="openai/gpt-oss-120b",
provider="cerebras",
allow_fallbacks=False,
reasoning_effort="low",
)
Free text
prediction = classifier.predict_text(
"Rome Against Rome is a 1964 sword-and-sandal film.",
mention="Rome Against Rome",
)
print(prediction.coarse_type) # CREATIVE_WORK
print(prediction.fine_type) # FILM
print(prediction.specific_type) # SWORD_AND_SANDAL_FILM
Only the supplied target mention is classified. The surrounding sentence is contextual evidence.
The same LLM call also returns backend-neutral candidate-retrieval metadata:
prediction = classifier.predict_text(
"Rmoe is the capital and largest city of Italy.",
mention="Rmoe",
)
print(prediction.retrieval_metadata.corrected_mention) # Rome
print(prediction.retrieval_metadata.context_keywords) # e.g. ("capital city", "Italy")
print(prediction.retrieval_metadata.wikipedia_urls)
# e.g. ("https://en.wikipedia.org/wiki/Rome",)
print(prediction.retrieval_metadata.dbpedia_urls)
# e.g. ("https://dbpedia.org/resource/Rome",)
print(prediction.high_level_reason)
Reference URLs are model predictions, not verified links. Invalid URL shapes,
Wikidata URLs, and QID-bearing values are removed locally, and every serialized
metadata object is explicitly marked unverified.
Structured input
prediction = classifier.predict_record(
{
"label": "Chrysler Cirrus",
"description": "mid-size four-door sedan model",
"manufacturer": "Chrysler",
}
)
print(prediction.coarse_type) # PRODUCT
print(prediction.fine_type) # VEHICLE_WEAPON_OR_EQUIPMENT_MODEL
print(prediction.subtype) # CAR_MODEL, when supported by the evidence
Existing QIDs, URLs, popularity, priors, and previous NER fields are not used as prediction evidence. Newly predicted reference URLs remain optional search hints and cannot determine the semantic type.
Tabular input
For one table cell, provide the column meaning and bounded row/column context:
prediction = classifier.predict_table_cell(
"Germany",
column_header="country name",
row_context={
"manufacturer": "Daimler AG",
"vehicle_model": "Chrysler Cirrus",
"assembly_location": "Sterling Heights, Michigan",
},
same_column_values=[
"Germany",
"United States",
"Canada",
],
table_name="vehicle_production.csv",
)
print(prediction.coarse_type) # LOCATION
print(prediction.fine_type) # COUNTRY_OR_SOVEREIGN_STATE
For multiple cells, use TableCellTask and predict_table_cells():
from wikidata_ner import TableCellTask
tasks = [
TableCellTask(
cell="Germany",
column_header="country name",
row_context={"manufacturer": "Daimler AG"},
same_column_values=["Germany", "United States", "Canada"],
),
TableCellTask(
cell="United States",
column_header="country name",
row_context={"manufacturer": "General Motors"},
same_column_values=["Germany", "United States", "Canada"],
),
]
predictions = classifier.predict_table_cells(tasks)
The production batch limit is 8 targets per physical request. Larger iterables are split into multiple requests automatically.
Prediction output
All prediction paths expose retrieval-oriented type fields such as:
coarse_typefine_typesubtypespecific_typeandspecific_typesretrieval_key,retrieval_path, andretrieval_tagsconfidenceabstainedandabstention_reason
LLM-backed mention predictions additionally expose:
retrieval_metadata, including corrected spelling, mention variants, disambiguating keywords, and optional Wikipedia/DBpedia reference URLshigh_level_reason, one explanation covering the type and metadata choices
Type and metadata confidence values are computed locally as evidence-derived posterior probabilities; numeric confidence self-ratings from the model are not used at any stage. This includes fine alternatives, subtypes, occupation or other facets, and retrieval metadata. Wikipedia and DBpedia URL hints are accepted only when the model explicitly returns them and they pass local domain, canonical-path, QID, and title checks; the library never constructs missing URLs. No other URL family or unique entity identifier is predicted.
Controlled type paths, keys, and tags are still validated and constructed
locally. Only the bounded retrieval_metadata hints are model-predicted.
to_candidate_retrieval_profile() provides a generic query contract that can be
adapted to Elasticsearch, a vector database, a knowledge-graph lookup, or another
retrieval system. Keep its roles separate:
- use
mention_queriesfor lexical candidate recall; - use
context_keywordsas soft disambiguation signals; - use hierarchy hints as confidence-aware filters or boosts;
- treat
reference_urlsas optional exact lookups that still require verification.
In particular, avoid concatenating every context keyword into the mention query: that can reward labels containing the context words rather than the entity named by the mention.
Release history
See CHANGELOG.md for version details.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file wikidata_ner_classifier-0.8.0.tar.gz.
File metadata
- Download URL: wikidata_ner_classifier-0.8.0.tar.gz
- Upload date:
- Size: 164.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.10.2 {"installer":{"name":"uv","version":"0.10.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1091835babc0da935c23a45f05f9cd63c416164faceb11cc699b8cb78bca40bf
|
|
| MD5 |
b8793f01090fdb68bb1eb90b2ebd225f
|
|
| BLAKE2b-256 |
f5e2d9fa98b398d23bc364c25f89e00a3f54c71dfec0c5a67f04b9484bfe2432
|
File details
Details for the file wikidata_ner_classifier-0.8.0-py3-none-any.whl.
File metadata
- Download URL: wikidata_ner_classifier-0.8.0-py3-none-any.whl
- Upload date:
- Size: 147.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.10.2 {"installer":{"name":"uv","version":"0.10.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3c0a1ea00833872e696a6ddfb68ab6b2222860ebdfceedd01493ec843aca78ba
|
|
| MD5 |
c3ebdb45bd14404bc4694a5f67286b1f
|
|
| BLAKE2b-256 |
14a02e888f3d975f9ea012d10c658299d6fea64e96ac459e7beb4edae07a2eec
|