Skip to main content

Wikidata NER Classifier 0.9.0

wikidata-ner-classifier predicts retrieval-oriented NER types in two ways:

  1. Wikidata items: deterministic prediction from P31/P279 token clues and an optional description.
  2. Input data: LLM prediction from a target mention and its context, supplied as free text or tabular data.

Both paths return types from the same hierarchy:

coarse_type -> fine_type -> subtype -> specific_type

The prediction can be used to narrow the candidate-retrieval space before entity linking. The library does not make the final identity decision. LLM predictions may include unverified Wikipedia and DBpedia URLs as search hints; a downstream linker must retrieve and verify them. Wikidata QIDs and URLs are deliberately excluded because they would be unique-identity guesses.

Installation

pip install wikidata-ner-classifier

1. Predict NER types for Wikidata items

Use WikidataNERClassifier when the input is already a Wikidata item and its P31/P279 type labels or aliases are available. Prediction is deterministic and does not require an LLM or network request.

from wikidata_ner import WikidataNERClassifier

classifier = WikidataNERClassifier()

prediction = classifier.predict(
    qid="Q3441181",
    types=[
        {"id": "Q11424", "name": "film"},
    ],
    description="1964 sword-and-sandal film directed by Giuseppe Vari",
)

print(prediction.coarse_type)  # CREATIVE_WORK
print(prediction.fine_type)  # FILM
print(prediction.specific_type)  # SWORD_AND_SANDAL_FILM
print(prediction.retrieval_key)
# CREATIVE_WORK/FILM/SWORD_AND_SANDAL_FILM

The P31/P279 token clues select the semantic branch. The description may refine the result, but it does not replace the type-token evidence.

A complete entity mapping can also be passed directly:

prediction = classifier.predict_entity(
    {
        "qid": "Q3441181",
        "types": [{"id": "Q11424", "name": "film"}],
        "description": "1964 sword-and-sandal film",
    }
)

For multiple Wikidata items, use predict_batch():

predictions = classifier.predict_batch(
    [
        {
            "qid": "Q3441181",
            "types": [{"name": "film"}],
            "description": "1964 sword-and-sandal film",
        },
        {
            "qid": "Q7259",
            "types": [{"name": "human"}],
            "description": "English mathematician and writer",
        },
    ]
)

2. Predict NER types for input data with an LLM

Use OpenRouterNERClassifier when the input is a mention whose type must be inferred from context. The context can be free text, a structured record, or a table cell.

Set an OpenRouter API key:

export OPENROUTER_API_KEY="..."

app_identifier= is sent as OpenRouter's HTTP-Referer attribution header. Its direct library default is https://github.com/roby-avo/ner-wikidata, so no environment variable is required and requests are not attributed to http://localhost/.

Create the classifier:

from wikidata_ner import OpenRouterNERClassifier

classifier = OpenRouterNERClassifier(
    model="openai/gpt-oss-120b",
    reasoning_effort="low",
    app_identifier="https://example.org/wikidata-ner-production",
)

For the current Cerebras-supported model IDs, the client automatically routes to Cerebras with fallbacks disabled when provider= is omitted. An explicit provider configuration remains an intentional override. Query generation also uses a bounded repair attempt when a weaker model returns a malformed response.

Free text

prediction = classifier.predict_text(
    "Rome Against Rome is a 1964 sword-and-sandal film.",
    mention="Rome Against Rome",
)

print(prediction.coarse_type)  # CREATIVE_WORK
print(prediction.fine_type)  # FILM
print(prediction.specific_type)  # SWORD_AND_SANDAL_FILM

Only the supplied target mention is classified. The surrounding sentence is contextual evidence.

The same LLM call also returns backend-neutral candidate-retrieval metadata:

prediction = classifier.predict_text(
    "Rmoe is the capital and largest city of Italy.",
    mention="Rmoe",
)

print(prediction.retrieval_metadata.corrected_mention)  # Rome
print(prediction.retrieval_metadata.surface_variants)  # local bounded repairs
print(prediction.retrieval_metadata.mention_query_signals)
# weighted lexical surfaces; the original mention is first
print(prediction.retrieval_metadata.context_keywords)  # e.g. ("capital city", "Italy")
print(prediction.retrieval_metadata.wikipedia_urls)
# e.g. ("https://en.wikipedia.org/wiki/Rome",)
print(prediction.retrieval_metadata.dbpedia_urls)
# e.g. ("https://dbpedia.org/resource/Rome",)
print(prediction.high_level_reason)

Reference URLs are model predictions, not verified links. Invalid URL shapes, Wikidata URLs, and QID-bearing values are removed locally, and every serialized metadata object is explicitly marked unverified.

Structured input

prediction = classifier.predict_record(
    {
        "label": "Chrysler Cirrus",
        "description": "mid-size four-door sedan model",
        "manufacturer": "Chrysler",
    }
)

print(prediction.coarse_type)  # PRODUCT
print(prediction.fine_type)  # VEHICLE_WEAPON_OR_EQUIPMENT_MODEL
print(prediction.subtype)  # CAR_MODEL, when supported by the evidence

Existing QIDs, URLs, popularity, priors, and previous NER fields are not used as prediction evidence. Newly predicted reference URLs remain optional search hints and cannot determine the semantic type.

Tabular input

For one table cell, provide the column meaning and bounded row/column context:

prediction = classifier.predict_table_cell(
    "Germany",
    column_header="country name",
    row_context={
        "manufacturer": "Daimler AG",
        "vehicle_model": "Chrysler Cirrus",
        "assembly_location": "Sterling Heights, Michigan",
    },
    same_column_values=[
        "Germany",
        "United States",
        "Canada",
    ],
    table_name="vehicle_production.csv",
)

print(prediction.coarse_type)  # LOCATION
print(prediction.fine_type)  # COUNTRY_OR_SOVEREIGN_STATE

For multiple cells, use TableCellTask and predict_table_cells():

from wikidata_ner import TableCellTask

tasks = [
    TableCellTask(
        cell="Germany",
        column_header="country name",
        row_context={"manufacturer": "Daimler AG"},
        same_column_values=["Germany", "United States", "Canada"],
    ),
    TableCellTask(
        cell="United States",
        column_header="country name",
        row_context={"manufacturer": "General Motors"},
        same_column_values=["Germany", "United States", "Canada"],
    ),
]

predictions = classifier.predict_table_cells(tasks)

The production batch limit is 8 targets per physical request. Larger iterables are split into multiple requests automatically.

Provider-supplied token and cost fields, along with request/model/provider identity and measured latency, are retained in prediction.usage when available.

Prediction output

All prediction paths expose retrieval-oriented type fields such as:

  • coarse_type
  • fine_type
  • subtype
  • specific_type and specific_types
  • retrieval_key, retrieval_path, and retrieval_tags
  • confidence
  • abstained and abstention_reason

LLM-backed mention predictions additionally expose:

  • retrieval_metadata, including corrected spelling, model mention variants, identity-preserving local surface_variants, weighted mention_query_signals, disambiguating keywords, and optional Wikipedia/DBpedia reference URLs
  • typed mention_variant_hints with transformation kind, source, and locally calibrated retrieval confidence; the existing mention_variants list remains available for backwards compatibility
  • high_level_reason, one explanation covering the type and metadata choices

Type confidence is computed locally with evidence-derived posterior odds; model numeric self-ratings are never used. Metadata signal values are separate empirical estimates of retrieval usefulness/non-harm, not entity-correctness probabilities. This includes explicit estimates for corrections, variants, keywords, and reference URLs. Wikipedia and DBpedia URL hints are accepted only when the model explicitly returns them and they pass local domain, canonical-path, QID, and title checks; the library never constructs missing URLs. No other URL family or unique entity identifier is predicted.

Controlled type paths, keys, and tags are still validated and constructed locally. Only the bounded retrieval_metadata hints are model-predicted.

to_candidate_retrieval_profile() provides a generic query contract that can be adapted to a lexical search service, a vector database, a knowledge-graph lookup, or another retrieval system. Keep its roles separate:

  • use mention_query_signals for weighted lexical candidate recall, always retaining the original mention as the strongest surface;
  • use context_keywords as soft disambiguation signals;
  • use hierarchy hints as confidence-aware filters or boosts;
  • treat reference_urls as optional exact lookups that still require verification.

In particular, avoid concatenating every context keyword into the mention query: that can reward labels containing the context words rather than the entity named by the mention.

Experimental candidate-retrieval query comparison

The package contains a backend-neutral experiment for comparing two candidate-query construction strategies while holding classifier metadata constant. It asks an injected candidate retrieval system for a schema, normalizes and fingerprints that response through an injected adapter, and persists artifacts/discovered_schema.json. A query-format-specific template query-template builder/compiler and optional query generator are then supplied by the caller. The artifacts are reused until the discovered schema fingerprint changes (or rediscovery is forced).

For a candidate API where the retrieval endpoint is already known, provide its complete URL and token. No backend-specific client is required:

import os

from wikidata_ner import HTTPAPICandidateRetrievalSystem

my_candidate_system = HTTPAPICandidateRetrievalSystem(
    url="https://candidate.example/api/search",
    api_token=os.environ["CANDIDATE_API_TOKEN"],
)

result = my_candidate_system.request_json(
    method="POST",
    payload={"mention": "Rome", "context": "capital of Italy", "top_k": 20},
)

When the retrieval endpoint itself is the source of candidate-field discovery, provide a discovery payload and an adapter that knows where candidate records live in the response. For example, an endpoint returning hits.hits[*]._source can be sampled with:

from pathlib import Path

from wikidata_ner import (
    CandidateSchemaDiscovery,
    SampledCandidateSchemaAdapter,
)

candidate_system = HTTPAPICandidateRetrievalSystem(
    url="https://candidate.example/api/search",
    api_token=os.environ["CANDIDATE_API_TOKEN"],
    discovery_payload={},
)

schema = CandidateSchemaDiscovery(
    system=candidate_system,
    adapter=SampledCandidateSchemaAdapter(
        records_path=("hits", "hits"),
        candidate_path=("_source",),
    ),
    artifact_path=Path("artifacts/discovered_schema.json"),
).load_or_discover(resource="documents")

This discovers fields and observed value types from sampled candidates. The response paths are configuration supplied by the integration; the experiment does not assume a particular retrieval backend.

For standard FastAPI/OpenAPI services, use base_url=...; the default schema discovery request is unauthenticated GET /openapi.json, matching the usual public FastAPI OpenAPI route. Set authenticate_schema_request=True when the schema endpoint is protected. Use OpenAPISchemaAdapter for standard OpenAPI documents. Custom OpenAPI URLs, token headers, and token prefixes are supported. A complete url=... can also be combined with openapi_url=... when the service exposes discovery at a separate URL.

For incremental setup, CandidateRetrievalSetupAgent performs the discovery request and asks the structured-completion client to identify and persist both a query contract and the candidate system's strict request-body JSON Schema. The contract records the discovered technology and its evidence, supported query types, field roles, and optimization guidance without embedding a backend-specific query implementation in the library or caller. That schema is tied to the discovery fingerprint. Runtime calls provide the complete contract, identified query schema, and input retrieval metadata together to the structured-completion client, then locally reject generated fields or values that do not conform to the schema. The candidate system is not rediscovered unless prepare(force=True) is requested:

import os
from pathlib import Path
from collections.abc import Mapping

from wikidata_ner import (
    CandidateRetrievalSetupAgent,
    HTTPAPICandidateRetrievalSystem,
    OpenRouterNERClassifier,
    metadata_from_prediction,
)


class RawDiscoveryAdapter:
    def normalize_schema(self, response, *, resource):
        if not isinstance(response, Mapping):
            raise TypeError("Discovery response must be a JSON object.")
        return dict(response)


candidate_system = HTTPAPICandidateRetrievalSystem(
    url=os.environ["CANDIDATE_RETRIEVAL_URL"],
    api_token=os.environ["CANDIDATE_API_TOKEN"],
    discovery_payload={},
)
classifier = OpenRouterNERClassifier(
    model="openai/gpt-oss-120b",
    api_key=os.environ["OPENROUTER_API_KEY"],
)

setup = CandidateRetrievalSetupAgent(
    resource="candidate-system",
    candidate_retrieval_system=candidate_system,
    candidate_schema_adapter=RawDiscoveryAdapter(),
    query_client=classifier.client,
    artifact_dir=Path("artifacts/candidate_setup"),
)

# One-time setup: discovery + AI-generated query contract.
profile = setup.prepare()
print(profile.query_schema)  # exact request-body schema identified at setup

prediction = classifier.predict_text(
    "Rmoe is the capital city of Italy.",
    mention="Rmoe",
)
metadata = metadata_from_prediction(prediction)

# Runtime: metadata + cached setup profile -> query.
query = setup.generate_query(metadata)
result = candidate_system.request_json(method="POST", payload=query)

Before technology-specific compilation, runtime generation derives a backend-neutral optimization plan from the metadata. It keeps all weighted identity surfaces as alternative lexical recall signals, applies context and unverified reference hints as optional boosts, and uses every predicted type as a confidence-scaled boost rather than a filter. The coverage-first plan requests up to 1,000 candidates and never lets context, type, popularity, prior, or an unverified URL determine eligibility. The discovered field roles and supported query types determine which parts can be compiled for the target system.

The setup profile, including its schema_fingerprint, query_contract, and query_schema, is persisted in artifacts/candidate_setup/discovery_profile.json. Legacy profiles without an identified query schema are analyzed again. Use setup.prepare(force=True) when the candidate-system contract may have changed.

from wikidata_ner import (
    CandidateRetrievalExperiment,
    LLMQueryGenerator,
    OpenRouterNERClassifier,
)

classifier = OpenRouterNERClassifier(model="openai/gpt-oss-120b")
experiment = CandidateRetrievalExperiment(
    resource="wikidata",
    candidate_retrieval_system=my_candidate_system,
    candidate_schema_adapter=my_schema_adapter,
    template_builder=my_query_template_builder,
    query_compiler=my_query_compiler,
    classifier=classifier,
    query_generator=LLMQueryGenerator(
        classifier.client,
        validator=my_query_validator,
    ),
)
comparison = experiment.build_queries(
    mention="Rmoe",
    context="Rmoe is the capital city of Italy.",
)

comparison.metadata is generated once and passed unchanged to both the deterministic schema-derived compiler (comparison.template_query) and the per-request LLM strategy (comparison.llm_query). The candidate system is queried only through CandidateRetrievalSystem; its discovery response is normalized by CandidateSchemaAdapter. Query syntax and validation remain outside this package.

For in-process candidate reranking, WikidataCandidateRanker.rank_many() accepts generic mappings and stops at max_candidates before scoring. The default and maximum are both 1,000, so an unbounded iterable cannot silently expand work. Medium-confidence hierarchy metadata uses coarse filtering with fine/specific boosts; low-confidence metadata is boost-only to protect recall.

When migrating serialized 0.8 source predictions, rebuild canonical path, key, tag, and level fields from the semantic type/facet inputs. Version 0.9 validates the unified path strictly and may reject the older three-level composite shape.

The live Cloudflare/CEA methodology, coverage comparisons, calibration caveats, and latency measurements for this release are recorded in benchmarks/RETRIEVAL_TUNING_2026-08-12.md.

Release history

See CHANGELOG.md for version details.

License

MIT

Release files for wikidata-ner-classifier 0.10.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for wikidata-ner-classifier 0.10.0
File Size Uploaded
wikidata_ner_classifier-0.10.0.tar.gz 231.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for wikidata-ner-classifier 0.10.0
File Interpreter ABI Platform
wikidata_ner_classifier-0.10.0-py3-none-any.whl Python 3 none any Details

Total release size: 407.6 kB

Release files / wikidata_ner_classifier-0.10.0.tar.gz

Download URL wikidata_ner_classifier-0.10.0.tar.gz
Size 231.5 kB
Tags Source
SHA-256 checksum
How to use checksums
9f64a86e271873065784e4219cc708ca4d95fba3f030d661864e4f105dda6689
BLAKE2b-256 checksum
How to use checksums
8a9650e8894039c0e958a8477f2678faa2cd688efa97ef17ddb6f2b4ec408f45
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.7

Release files / wikidata_ner_classifier-0.10.0-py3-none-any.whl

Download URL wikidata_ner_classifier-0.10.0-py3-none-any.whl
Size 176.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6cc1b723d6104ac492834009cfd070697ccd62677f47bef23f896bd236a42e1d
BLAKE2b-256 checksum
How to use checksums
181ab0c9df4b739f64705258d2c2f46b41161e373aa8504eec7850e3b9b45fb0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.7

Release history Release notifications | RSS feed

0.10.1

2 release files

This release

0.10.0 This release

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page