Wikidata NER Classifier 0.9.0
wikidata-ner-classifier predicts retrieval-oriented NER types in two ways:
- Wikidata items: deterministic prediction from P31/P279 token clues and an optional description.
- Input data: LLM prediction from a target mention and its context, supplied as free text or tabular data.
Both paths return types from the same hierarchy:
coarse_type -> fine_type -> subtype -> specific_type
The prediction can be used to narrow the candidate-retrieval space before entity linking. The library does not make the final identity decision. LLM predictions may include unverified Wikipedia and DBpedia URLs as search hints; a downstream linker must retrieve and verify them. Wikidata QIDs and URLs are deliberately excluded because they would be unique-identity guesses.
Installation
pip install wikidata-ner-classifier
1. Predict NER types for Wikidata items
Use WikidataNERClassifier when the input is already a Wikidata item and its
P31/P279 type labels or aliases are available. Prediction is deterministic and
does not require an LLM or network request.
from wikidata_ner import WikidataNERClassifier
classifier = WikidataNERClassifier()
prediction = classifier.predict(
qid="Q3441181",
types=[
{"id": "Q11424", "name": "film"},
],
description="1964 sword-and-sandal film directed by Giuseppe Vari",
)
print(prediction.coarse_type) # CREATIVE_WORK
print(prediction.fine_type) # FILM
print(prediction.specific_type) # SWORD_AND_SANDAL_FILM
print(prediction.retrieval_key)
# CREATIVE_WORK/FILM/SWORD_AND_SANDAL_FILM
The P31/P279 token clues select the semantic branch. The description may refine the result, but it does not replace the type-token evidence.
A complete entity mapping can also be passed directly:
prediction = classifier.predict_entity(
{
"qid": "Q3441181",
"types": [{"id": "Q11424", "name": "film"}],
"description": "1964 sword-and-sandal film",
}
)
For multiple Wikidata items, use predict_batch():
predictions = classifier.predict_batch(
[
{
"qid": "Q3441181",
"types": [{"name": "film"}],
"description": "1964 sword-and-sandal film",
},
{
"qid": "Q7259",
"types": [{"name": "human"}],
"description": "English mathematician and writer",
},
]
)
2. Predict NER types for input data with an LLM
Use OpenRouterNERClassifier when the input is a mention whose type must be
inferred from context. The context can be free text, a structured record, or a
table cell.
Set an OpenRouter API key:
export OPENROUTER_API_KEY="..."
app_identifier= is sent as OpenRouter's HTTP-Referer attribution header. Its
direct library default is https://github.com/roby-avo/ner-wikidata, so no
environment variable is required and requests are not attributed to
http://localhost/.
Create the classifier:
from wikidata_ner import OpenRouterNERClassifier
classifier = OpenRouterNERClassifier(
model="openai/gpt-oss-120b",
reasoning_effort="low",
app_identifier="https://example.org/wikidata-ner-production",
)
For the current Cerebras-supported model IDs, the client automatically routes
to Cerebras with fallbacks disabled when provider= is omitted. An explicit
provider configuration remains an intentional override. Query generation also
uses a bounded repair attempt when a weaker model returns a malformed response.
Free text
prediction = classifier.predict_text(
"Rome Against Rome is a 1964 sword-and-sandal film.",
mention="Rome Against Rome",
)
print(prediction.coarse_type) # CREATIVE_WORK
print(prediction.fine_type) # FILM
print(prediction.specific_type) # SWORD_AND_SANDAL_FILM
Only the supplied target mention is classified. The surrounding sentence is contextual evidence.
The same LLM call also returns backend-neutral candidate-retrieval metadata:
prediction = classifier.predict_text(
"Rmoe is the capital and largest city of Italy.",
mention="Rmoe",
)
print(prediction.retrieval_metadata.corrected_mention) # Rome
print(prediction.retrieval_metadata.surface_variants) # local bounded repairs
print(prediction.retrieval_metadata.mention_query_signals)
# weighted lexical surfaces; the original mention is first
print(prediction.retrieval_metadata.context_keywords) # e.g. ("capital city", "Italy")
print(prediction.retrieval_metadata.wikipedia_urls)
# e.g. ("https://en.wikipedia.org/wiki/Rome",)
print(prediction.retrieval_metadata.dbpedia_urls)
# e.g. ("https://dbpedia.org/resource/Rome",)
print(prediction.high_level_reason)
Reference URLs are model predictions, not verified links. Invalid URL shapes,
Wikidata URLs, and QID-bearing values are removed locally, and every serialized
metadata object is explicitly marked unverified.
Structured input
prediction = classifier.predict_record(
{
"label": "Chrysler Cirrus",
"description": "mid-size four-door sedan model",
"manufacturer": "Chrysler",
}
)
print(prediction.coarse_type) # PRODUCT
print(prediction.fine_type) # VEHICLE_WEAPON_OR_EQUIPMENT_MODEL
print(prediction.subtype) # CAR_MODEL, when supported by the evidence
Existing QIDs, URLs, popularity, priors, and previous NER fields are not used as prediction evidence. Newly predicted reference URLs remain optional search hints and cannot determine the semantic type.
Tabular input
For one table cell, provide the column meaning and bounded row/column context:
prediction = classifier.predict_table_cell(
"Germany",
column_header="country name",
row_context={
"manufacturer": "Daimler AG",
"vehicle_model": "Chrysler Cirrus",
"assembly_location": "Sterling Heights, Michigan",
},
same_column_values=[
"Germany",
"United States",
"Canada",
],
table_name="vehicle_production.csv",
)
print(prediction.coarse_type) # LOCATION
print(prediction.fine_type) # COUNTRY_OR_SOVEREIGN_STATE
For multiple cells, use TableCellTask and predict_table_cells():
from wikidata_ner import TableCellTask
tasks = [
TableCellTask(
cell="Germany",
column_header="country name",
row_context={"manufacturer": "Daimler AG"},
same_column_values=["Germany", "United States", "Canada"],
),
TableCellTask(
cell="United States",
column_header="country name",
row_context={"manufacturer": "General Motors"},
same_column_values=["Germany", "United States", "Canada"],
),
]
predictions = classifier.predict_table_cells(tasks)
The production batch limit is 8 targets per physical request. Larger iterables are split into multiple requests automatically.
Provider-supplied token and cost fields, along with request/model/provider
identity and measured latency, are retained in prediction.usage when
available.
Prediction output
All prediction paths expose retrieval-oriented type fields such as:
coarse_typefine_typesubtypespecific_typeandspecific_typesretrieval_key,retrieval_path, andretrieval_tagsconfidenceabstainedandabstention_reason
LLM-backed mention predictions additionally expose:
retrieval_metadata, including corrected spelling, model mention variants, identity-preserving localsurface_variants, weightedmention_query_signals, disambiguating keywords, and optional Wikipedia/DBpedia reference URLs- typed
mention_variant_hintswith transformation kind, source, and locally calibrated retrieval confidence; the existingmention_variantslist remains available for backwards compatibility high_level_reason, one explanation covering the type and metadata choices
Type confidence is computed locally with evidence-derived posterior odds; model numeric self-ratings are never used. Metadata signal values are separate empirical estimates of retrieval usefulness/non-harm, not entity-correctness probabilities. This includes explicit estimates for corrections, variants, keywords, and reference URLs. Wikipedia and DBpedia URL hints are accepted only when the model explicitly returns them and they pass local domain, canonical-path, QID, and title checks; the library never constructs missing URLs. No other URL family or unique entity identifier is predicted.
Controlled type paths, keys, and tags are still validated and constructed
locally. Only the bounded retrieval_metadata hints are model-predicted.
to_candidate_retrieval_profile() provides a generic query contract that can be
adapted to a lexical search service, a vector database, a knowledge-graph lookup,
or another retrieval system. Keep its roles separate:
- use
mention_query_signalsfor weighted lexical candidate recall, always retaining the original mention as the strongest surface; - use
context_keywordsas soft disambiguation signals; - use hierarchy hints as confidence-aware filters or boosts;
- treat
reference_urlsas optional exact lookups that still require verification.
In particular, avoid concatenating every context keyword into the mention query: that can reward labels containing the context words rather than the entity named by the mention.
For in-process candidate reranking, WikidataCandidateRanker.rank_many() accepts
generic mappings and stops at max_candidates before scoring. The default and
maximum are both 1,000, so an unbounded iterable cannot silently expand work.
Medium-confidence hierarchy metadata uses coarse filtering with fine/specific
boosts; low-confidence metadata is boost-only to protect recall.
When migrating serialized 0.8 source predictions, rebuild canonical path, key, tag, and level fields from the semantic type/facet inputs. Version 0.9 validates the unified path strictly and may reject the older three-level composite shape.
The live Cloudflare/CEA methodology, coverage comparisons, calibration caveats,
and latency measurements for this release are recorded in
benchmarks/RETRIEVAL_TUNING_2026-08-12.md.
Release history
See CHANGELOG.md for version details.
License
MIT
Release files for wikidata-ner-classifier 0.10.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| wikidata_ner_classifier-0.10.1.tar.gz | 209.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| wikidata_ner_classifier-0.10.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 370.6 kB
Release files / wikidata_ner_classifier-0.10.1.tar.gz
| Download URL | wikidata_ner_classifier-0.10.1.tar.gz |
|---|---|
| Size | 209.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6ba58312d312129fc88c0dba55ab9f91da61efd767beb4f55cf826b4cb897c83
|
|
BLAKE2b-256 checksum How to use checksums |
fad2066e5b3835da13a83b9afee6858ba907796be8340d08582e4929a2e2a849
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|
Release files / wikidata_ner_classifier-0.10.1-py3-none-any.whl
| Download URL | wikidata_ner_classifier-0.10.1-py3-none-any.whl |
|---|---|
| Size | 160.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
09ed9702dcfda2ac23108ffce39b2a30868820295772513e903bc373c9d1feff
|
|
BLAKE2b-256 checksum How to use checksums |
c4434d9d4dc5b2c16cbcbd87f9ec2561d1883b0c9b03fc7130e8a9524bf11e6d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|