Wikidata NER Classifier 0.9.0
wikidata-ner-classifier predicts retrieval-oriented NER types in two ways:
- Wikidata items: deterministic prediction from P31/P279 token clues and an optional description.
- Input data: LLM prediction from a target mention and its context, supplied as free text or tabular data.
Both paths return types from the same hierarchy:
coarse_type -> fine_type -> subtype -> specific_type
The prediction can be used to narrow the candidate-retrieval space before entity linking. The library does not make the final identity decision. LLM predictions may include unverified Wikipedia and DBpedia URLs as search hints; a downstream linker must retrieve and verify them. Wikidata QIDs and URLs are deliberately excluded because they would be unique-identity guesses.
Installation
pip install wikidata-ner-classifier
1. Predict NER types for Wikidata items
Use WikidataNERClassifier when the input is already a Wikidata item and its
P31/P279 type labels or aliases are available. Prediction is deterministic and
does not require an LLM or network request.
from wikidata_ner import WikidataNERClassifier
classifier = WikidataNERClassifier()
prediction = classifier.predict(
qid="Q3441181",
types=[
{"id": "Q11424", "name": "film"},
],
description="1964 sword-and-sandal film directed by Giuseppe Vari",
)
print(prediction.coarse_type) # CREATIVE_WORK
print(prediction.fine_type) # FILM
print(prediction.specific_type) # SWORD_AND_SANDAL_FILM
print(prediction.retrieval_key)
# CREATIVE_WORK/FILM/SWORD_AND_SANDAL_FILM
The P31/P279 token clues select the semantic branch. The description may refine the result, but it does not replace the type-token evidence.
A complete entity mapping can also be passed directly:
prediction = classifier.predict_entity(
{
"qid": "Q3441181",
"types": [{"id": "Q11424", "name": "film"}],
"description": "1964 sword-and-sandal film",
}
)
For multiple Wikidata items, use predict_batch():
predictions = classifier.predict_batch(
[
{
"qid": "Q3441181",
"types": [{"name": "film"}],
"description": "1964 sword-and-sandal film",
},
{
"qid": "Q7259",
"types": [{"name": "human"}],
"description": "English mathematician and writer",
},
]
)
2. Predict NER types for input data with an LLM
Use OpenRouterNERClassifier when the input is a mention whose type must be
inferred from context. The context can be free text, a structured record, or a
table cell.
Set an OpenRouter API key:
export OPENROUTER_API_KEY="..."
app_identifier= is sent as OpenRouter's HTTP-Referer attribution header. Its
direct library default is https://github.com/roby-avo/ner-wikidata, so no
environment variable is required and requests are not attributed to
http://localhost/.
Create the classifier:
from wikidata_ner import OpenRouterNERClassifier
classifier = OpenRouterNERClassifier(
model="openai/gpt-oss-120b",
reasoning_effort="low",
app_identifier="https://example.org/wikidata-ner-production",
)
For the current Cerebras-supported model IDs, the client automatically routes
to Cerebras with fallbacks disabled when provider= is omitted. An explicit
provider configuration remains an intentional override. Query generation also
uses a bounded repair attempt when a weaker model returns a malformed response.
Free text
prediction = classifier.predict_text(
"Rome Against Rome is a 1964 sword-and-sandal film.",
mention="Rome Against Rome",
)
print(prediction.coarse_type) # CREATIVE_WORK
print(prediction.fine_type) # FILM
print(prediction.specific_type) # SWORD_AND_SANDAL_FILM
Only the supplied target mention is classified. The surrounding sentence is contextual evidence.
The same LLM call also returns backend-neutral candidate-retrieval metadata:
prediction = classifier.predict_text(
"Rmoe is the capital and largest city of Italy.",
mention="Rmoe",
)
print(prediction.retrieval_metadata.corrected_mention) # Rome
print(prediction.retrieval_metadata.surface_variants) # local bounded repairs
print(prediction.retrieval_metadata.mention_query_signals)
# weighted lexical surfaces; the original mention is first
print(prediction.retrieval_metadata.context_keywords) # e.g. ("capital city", "Italy")
print(prediction.retrieval_metadata.wikipedia_urls)
# e.g. ("https://en.wikipedia.org/wiki/Rome",)
print(prediction.retrieval_metadata.dbpedia_urls)
# e.g. ("https://dbpedia.org/resource/Rome",)
print(prediction.high_level_reason)
Reference URLs are model predictions, not verified links. Invalid URL shapes,
Wikidata URLs, and QID-bearing values are removed locally, and every serialized
metadata object is explicitly marked unverified.
Structured input
prediction = classifier.predict_record(
{
"label": "Chrysler Cirrus",
"description": "mid-size four-door sedan model",
"manufacturer": "Chrysler",
}
)
print(prediction.coarse_type) # PRODUCT
print(prediction.fine_type) # VEHICLE_WEAPON_OR_EQUIPMENT_MODEL
print(prediction.subtype) # CAR_MODEL, when supported by the evidence
Existing QIDs, URLs, popularity, priors, and previous NER fields are not used as prediction evidence. Newly predicted reference URLs remain optional search hints and cannot determine the semantic type.
Tabular input
For one table cell, provide the column meaning and bounded row/column context:
prediction = classifier.predict_table_cell(
"Germany",
column_header="country name",
row_context={
"manufacturer": "Daimler AG",
"vehicle_model": "Chrysler Cirrus",
"assembly_location": "Sterling Heights, Michigan",
},
same_column_values=[
"Germany",
"United States",
"Canada",
],
table_name="vehicle_production.csv",
)
print(prediction.coarse_type) # LOCATION
print(prediction.fine_type) # COUNTRY_OR_SOVEREIGN_STATE
For multiple cells, use TableCellTask and predict_table_cells():
from wikidata_ner import TableCellTask
tasks = [
TableCellTask(
cell="Germany",
column_header="country name",
row_context={"manufacturer": "Daimler AG"},
same_column_values=["Germany", "United States", "Canada"],
),
TableCellTask(
cell="United States",
column_header="country name",
row_context={"manufacturer": "General Motors"},
same_column_values=["Germany", "United States", "Canada"],
),
]
predictions = classifier.predict_table_cells(tasks)
The production batch limit is 8 targets per physical request. Larger iterables are split into multiple requests automatically.
Provider-supplied token and cost fields, along with request/model/provider
identity and measured latency, are retained in prediction.usage when
available.
Prediction output
All prediction paths expose retrieval-oriented type fields such as:
coarse_typefine_typesubtypespecific_typeandspecific_typesretrieval_key,retrieval_path, andretrieval_tagsconfidenceabstainedandabstention_reason
LLM-backed mention predictions additionally expose:
retrieval_metadata, including corrected spelling, model mention variants, identity-preserving localsurface_variants, weightedmention_query_signals, disambiguating keywords, and optional Wikipedia/DBpedia reference URLs- typed
mention_variant_hintswith transformation kind, source, and locally calibrated retrieval confidence; the existingmention_variantslist remains available for backwards compatibility high_level_reason, one explanation covering the type and metadata choices
Type confidence is computed locally with evidence-derived posterior odds; model numeric self-ratings are never used. Metadata signal values are separate empirical estimates of retrieval usefulness/non-harm, not entity-correctness probabilities. This includes explicit estimates for corrections, variants, keywords, and reference URLs. Wikipedia and DBpedia URL hints are accepted only when the model explicitly returns them and they pass local domain, canonical-path, QID, and title checks; the library never constructs missing URLs. No other URL family or unique entity identifier is predicted.
Controlled type paths, keys, and tags are still validated and constructed
locally. Only the bounded retrieval_metadata hints are model-predicted.
to_candidate_retrieval_profile() provides a generic query contract that can be
adapted to a lexical search service, a vector database, a knowledge-graph lookup,
or another retrieval system. Keep its roles separate:
- use
mention_query_signalsfor weighted lexical candidate recall, always retaining the original mention as the strongest surface; - use
context_keywordsas soft disambiguation signals; - use hierarchy hints as confidence-aware filters or boosts;
- treat
reference_urlsas optional exact lookups that still require verification.
In particular, avoid concatenating every context keyword into the mention query: that can reward labels containing the context words rather than the entity named by the mention.
Experimental candidate-retrieval query comparison
The package contains a backend-neutral experiment for comparing two
candidate-query construction strategies while holding classifier metadata
constant. It asks an injected candidate retrieval system for a schema,
normalizes and fingerprints that response through an injected adapter, and
persists artifacts/discovered_schema.json. A query-format-specific template
query-template builder/compiler and optional query generator are then supplied
by the caller.
The artifacts are reused until the discovered schema fingerprint changes (or
rediscovery is forced).
For a candidate API where the retrieval endpoint is already known, provide its complete URL and token. No backend-specific client is required:
import os
from wikidata_ner import HTTPAPICandidateRetrievalSystem
my_candidate_system = HTTPAPICandidateRetrievalSystem(
url="https://candidate.example/api/search",
api_token=os.environ["CANDIDATE_API_TOKEN"],
)
result = my_candidate_system.request_json(
method="POST",
payload={"mention": "Rome", "context": "capital of Italy", "top_k": 20},
)
When the retrieval endpoint itself is the source of candidate-field discovery,
provide a discovery payload and an adapter that knows where candidate records
live in the response. For example, an endpoint returning
hits.hits[*]._source can be sampled with:
from pathlib import Path
from wikidata_ner import (
CandidateSchemaDiscovery,
SampledCandidateSchemaAdapter,
)
candidate_system = HTTPAPICandidateRetrievalSystem(
url="https://candidate.example/api/search",
api_token=os.environ["CANDIDATE_API_TOKEN"],
discovery_payload={},
)
schema = CandidateSchemaDiscovery(
system=candidate_system,
adapter=SampledCandidateSchemaAdapter(
records_path=("hits", "hits"),
candidate_path=("_source",),
),
artifact_path=Path("artifacts/discovered_schema.json"),
).load_or_discover(resource="documents")
This discovers fields and observed value types from sampled candidates. The response paths are configuration supplied by the integration; the experiment does not assume a particular retrieval backend.
For standard FastAPI/OpenAPI services, use base_url=...; the default schema
discovery request is unauthenticated GET /openapi.json, matching the usual
public FastAPI OpenAPI route. Set authenticate_schema_request=True when the
schema endpoint is protected. Use OpenAPISchemaAdapter for standard OpenAPI
documents. Custom OpenAPI URLs, token headers, and token prefixes are
supported. A complete url=... can also be combined with openapi_url=...
when the service exposes discovery at a separate URL.
For incremental setup, CandidateRetrievalSetupAgent performs the discovery
request and asks the structured-completion client to identify and persist both
a query contract and the candidate system's strict request-body JSON Schema.
The contract records the discovered technology and its evidence, supported
query types, field roles, and optimization guidance without embedding a
backend-specific query implementation in the library or caller. That schema is
tied to the discovery fingerprint. Runtime calls provide the complete contract,
identified query schema, and input retrieval metadata together to the
structured-completion client, then locally reject generated fields or values
that do not conform to the schema. The candidate system is not rediscovered
unless prepare(force=True) is requested:
import os
from pathlib import Path
from collections.abc import Mapping
from wikidata_ner import (
CandidateRetrievalSetupAgent,
HTTPAPICandidateRetrievalSystem,
OpenRouterNERClassifier,
metadata_from_prediction,
)
class RawDiscoveryAdapter:
def normalize_schema(self, response, *, resource):
if not isinstance(response, Mapping):
raise TypeError("Discovery response must be a JSON object.")
return dict(response)
candidate_system = HTTPAPICandidateRetrievalSystem(
url=os.environ["CANDIDATE_RETRIEVAL_URL"],
api_token=os.environ["CANDIDATE_API_TOKEN"],
discovery_payload={},
)
classifier = OpenRouterNERClassifier(
model="openai/gpt-oss-120b",
api_key=os.environ["OPENROUTER_API_KEY"],
)
setup = CandidateRetrievalSetupAgent(
resource="candidate-system",
candidate_retrieval_system=candidate_system,
candidate_schema_adapter=RawDiscoveryAdapter(),
query_client=classifier.client,
artifact_dir=Path("artifacts/candidate_setup"),
)
# One-time setup: discovery + AI-generated query contract.
profile = setup.prepare()
print(profile.query_schema) # exact request-body schema identified at setup
prediction = classifier.predict_text(
"Rmoe is the capital city of Italy.",
mention="Rmoe",
)
metadata = metadata_from_prediction(prediction)
# Runtime: metadata + cached setup profile -> query.
query = setup.generate_query(metadata)
result = candidate_system.request_json(method="POST", payload=query)
Before technology-specific compilation, runtime generation derives a backend-neutral optimization plan from the metadata. It keeps all weighted identity surfaces as alternative lexical recall signals, applies context and unverified reference hints as optional boosts, and uses every predicted type as a confidence-scaled boost rather than a filter. The coverage-first plan requests up to 1,000 candidates and never lets context, type, popularity, prior, or an unverified URL determine eligibility. The discovered field roles and supported query types determine which parts can be compiled for the target system.
The setup profile, including its schema_fingerprint, query_contract, and
query_schema, is persisted in
artifacts/candidate_setup/discovery_profile.json. Legacy profiles without an
identified query schema are analyzed again. Use setup.prepare(force=True)
when the candidate-system contract may have changed.
from wikidata_ner import (
CandidateRetrievalExperiment,
LLMQueryGenerator,
OpenRouterNERClassifier,
)
classifier = OpenRouterNERClassifier(model="openai/gpt-oss-120b")
experiment = CandidateRetrievalExperiment(
resource="wikidata",
candidate_retrieval_system=my_candidate_system,
candidate_schema_adapter=my_schema_adapter,
template_builder=my_query_template_builder,
query_compiler=my_query_compiler,
classifier=classifier,
query_generator=LLMQueryGenerator(
classifier.client,
validator=my_query_validator,
),
)
comparison = experiment.build_queries(
mention="Rmoe",
context="Rmoe is the capital city of Italy.",
)
comparison.metadata is generated once and passed unchanged to both the
deterministic schema-derived compiler (comparison.template_query) and the
per-request LLM strategy (comparison.llm_query). The candidate system is
queried only through CandidateRetrievalSystem; its discovery response is
normalized by CandidateSchemaAdapter. Query syntax and validation remain
outside this package.
For in-process candidate reranking, WikidataCandidateRanker.rank_many() accepts
generic mappings and stops at max_candidates before scoring. The default and
maximum are both 1,000, so an unbounded iterable cannot silently expand work.
Medium-confidence hierarchy metadata uses coarse filtering with fine/specific
boosts; low-confidence metadata is boost-only to protect recall.
When migrating serialized 0.8 source predictions, rebuild canonical path, key, tag, and level fields from the semantic type/facet inputs. Version 0.9 validates the unified path strictly and may reject the older three-level composite shape.
The live Cloudflare/CEA methodology, coverage comparisons, calibration caveats,
and latency measurements for this release are recorded in
benchmarks/RETRIEVAL_TUNING_2026-08-12.md.
Release history
See CHANGELOG.md for version details.
License
MIT
Release files for wikidata-ner-classifier 0.10.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| wikidata_ner_classifier-0.10.0.tar.gz | 231.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| wikidata_ner_classifier-0.10.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 407.6 kB
Release files / wikidata_ner_classifier-0.10.0.tar.gz
| Download URL | wikidata_ner_classifier-0.10.0.tar.gz |
|---|---|
| Size | 231.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9f64a86e271873065784e4219cc708ca4d95fba3f030d661864e4f105dda6689
|
|
BLAKE2b-256 checksum How to use checksums |
8a9650e8894039c0e958a8477f2678faa2cd688efa97ef17ddb6f2b4ec408f45
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|
Release files / wikidata_ner_classifier-0.10.0-py3-none-any.whl
| Download URL | wikidata_ner_classifier-0.10.0-py3-none-any.whl |
|---|---|
| Size | 176.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6cc1b723d6104ac492834009cfd070697ccd62677f47bef23f896bd236a42e1d
|
|
BLAKE2b-256 checksum How to use checksums |
181ab0c9df4b739f64705258d2c2f46b41161e373aa8504eec7850e3b9b45fb0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|