Skip to main content

gdelt-regkg

CI PyPI Python License

Package for reconstructing GDELT GKG files from article text and metadata.

gdelt-regkg converts prepared article bodies and supplied metadata into GKG records. It can be used for reconstructing gaps in GKG coverage, processing your own news datasets, or extract individual fields for analysis. To make it work you just need article text and some metadata, while the library extracts and formats the fields automatically. You can generate full GKG entry or extract data only for specific fields.

The row structure, metadata formatting and field encodings are documented in GDELT Global Knowledge Graph Codebook. Extraction from text approximates GDELT's processing, so themes, counts, entities, tone and other inferred values can differ from native GKG results.

See the field documentation and measured agreement, resource and benchmark tools, and changelog.

Installation

Requires Python 3.12 or newer. Install the library and its English model, then prepare the extraction dictionaries:

python -m pip install gdelt-regkg
python -m spacy download en_core_web_sm
gdelt-regkg-build-resources --sources-dir downloads/resources --download-sources --output local-resources

Installation and setup explains what is downloaded, how to run your first extraction, and how to add larger models, GCAM dictionaries or translation.

Usage

Extract individual fields

Each field has a standalone function, so you do not need to generate a full GKG row.

from gdelt_regkg.fields import extract_counts, extract_themes, extract_tone

text = "Flooding killed 12 people and displaced 300 residents."

counts = extract_counts(text)
# CRISISLEX_T03_DEAD#12#people#0######;KILL#12#people#0######;AFFECT#300#residents#0######;CRISISLEX_CRISISLEXREC#300#residents#0######;CRISISLEX_T09_DISPLACEDRELOCATEDEVACUATED#300#residents#0######;DISPLACED#300#residents#0######;
themes = extract_themes(text)
# CRISISLEX_C06_WATER_SANITATION;NATURAL_DISASTER_FLOODING;CRISISLEX_T03_DEAD;KILL;AFFECT;CRISISLEX_CRISISLEXREC;CRISISLEX_T09_DISPLACEDRELOCATEDEVACUATED;DISPLACED;MANMADE_DISASTER_IMPLIED;
tone = extract_tone(text)
# -12.5,0,12.5,12.5,12.5,0,8

Functions return GKG-formatted cell strings and work with pandas, for example df["themes"] = df["text"].map(extract_themes). pandas package is optional and must be installed separately.

Tone, themes, geographic references and GCAM require locally prepared profiles. The default directory is local-resources/; select another with resources_dir= in Python or --resources-dir in the CLI. Explicit resource objects override directory defaults. Counts and name rules are available directly from the package.

Generate and save a row

from gdelt_regkg import SpacyNER, generate_gkg, write_gkg

ner = SpacyNER.from_model("en_core_web_sm", name_policy="surface")  # Load once and reuse.

record = generate_gkg(
    "Volunteers rescued 12 people in Paris, France.",
    identifier="https://example.org/story",
    published_at="20261001090023",  # Known precise publication time, stored in extra XML.
    batch_time="20261001091500",
    sequence=1,
    ner=ner,
    resources_dir="local-resources",
    sharing_image="",
    related_images=[],
    social_image_embeds=[],
    social_video_embeds=[],
    extras={},
    strict=True,
)

with open("example.gkg.csv", "w", encoding="utf-8", newline="") as stream:
    write_gkg([record], stream)

for column in ("V1COUNTS", "V1LOCATIONS", "V2.1AMOUNTS"):
    print(f"{column}: {record[column]}")
V1COUNTS: CRISISLEX_T08_MISSINGFOUNDTRAPPEDPEOPLE#12#people#4#Paris, France (General), France#FR#FR00#48.8667#2.33333#-1456928;
V1LOCATIONS: 4#Paris, France (General), France#FR#FR00#48.8667#2.33333#-1456928;1#France#FR#FR#46#2#FR
V2.1AMOUNTS: 12,people,19

These are three cells from the generated record. example.gkg.csv contains the complete 27-column GKG entry, including themes, tone, GCAM and enhanced fields. The generator resolves locations from NER and adds the nearest resolved location to each count.

The generator returns a dictionary of 27 fields. This example explicitly supplies a recognizer and empty media/extras metadata, so strict=True verifies a complete row. Missing supplied metadata or model-dependent fields without ner remain None and produce a warning. allow_incomplete=True exports those unavailable values as blank cells. Output is headerless and tab-delimited, even when the filename ends in .csv (GDELT specifics).

For non-English articles and the full GCAM profile, use the complete pipeline example in the setup guide.

Field implementation map

Columns appear in output order. Model-dependent fields require an installed recognizer; missing supplied metadata remains None.

Column Standalone function Status
GKGRECORDID format_record_id Supplied metadata
V2.1DATE format_date Batch timestamp
V2SOURCECOLLECTIONIDENTIFIER format_source_collection Supplied metadata
V2SOURCECOMMONNAME source_common_name Explicit label or hostname
V2DOCUMENTIDENTIFIER document_identifier Supplied metadata
V1COUNTS extract_counts Experimental approximation
V2.1COUNTS extract_enhanced_counts Same counts with number offsets
V1THEMES extract_themes Rule-based approximation
V2ENHANCEDTHEMES extract_enhanced_themes Same theme mentions with phrase offsets
V1LOCATIONS extract_locations Offline reference gazetteer and NER approximation
V2ENHANCEDLOCATIONS extract_enhanced_locations Same resolved places with ADM2 and offsets
V1PERSONS extract_persons spaCy NER approximation
V2ENHANCEDPERSONS extract_enhanced_persons Same person mentions with offsets
V1ORGANIZATIONS extract_organizations spaCy NER approximation
V2ENHANCEDORGANIZATIONS extract_enhanced_organizations Same organization mentions with offsets
V1.5TONE extract_tone Benchmark-refined dictionary approximation
V2.1ENHANCEDDATES extract_enhanced_dates Explicit English dates with resolutions 1–4
V2GCAM extract_gcam Dictionary category counts and weighted scores; coverage depends on local dictionaries
V2.1SHARINGIMAGE format_sharing_image Supplied image URL
V2.1RELATEDIMAGES format_related_images Supplied image URLs
V2.1SOCIALIMAGEEMBEDS format_social_image_embeds Supplied social-post URLs
V2.1SOCIALVIDEOEMBEDS format_social_video_embeds Supplied video URLs
V2.1QUOTATIONS extract_quotations Quoted spans and reporting-verb approximation
V2.1ALLNAMES extract_all_names Broad named entities and proper-noun spans
V2.1AMOUNTS extract_amounts Numeric/textual quantities with bounded objects
V2.1TRANSLATIONINFO format_translation_info Supplied metadata
V2EXTRASXML extract_extras_xml Custom extension data, including supplied metadata

See the field documentation for implementation status, research, and usage for each field.

Contributing

Use the GitHub issue forms to report bugs, extraction differences or feature requests. Contributing describes local checks, pull requests and evaluation on data separate from rule development. Please follow the code of conduct. Report vulnerabilities privately as described in the security policy.

Disclaimer

This is an independent implementation, unaffiliated with and not endorsed by the GDELT Project. It produces approximate GKG fields, not guaranteed copies of GDELT's outputs. Results depend on article text, translation, models and dictionaries; validate them for your intended analysis. Differences from native GKG can also reflect different input bodies.

The code and original rules use the MIT license and are provided without warranty. MIT does not grant rights to external dictionaries, models, article bodies or generated profiles derived from them; their source terms still apply.

Metadata

Release files for gdelt-regkg 1.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gdelt-regkg 1.0.2
File Size Uploaded
gdelt_regkg-1.0.2.tar.gz 438.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gdelt-regkg 1.0.2
File Interpreter ABI Platform
gdelt_regkg-1.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 808.3 kB

Release files / gdelt_regkg-1.0.2.tar.gz

Download URL gdelt_regkg-1.0.2.tar.gz
Size 438.9 kB
Tags Source
SHA-256 checksum
How to use checksums
8ffad64f1108ec20d6d879431c212bc67bae1e18773712fca8da22b36c194b7b
BLAKE2b-256 checksum
How to use checksums
aa71b5a7fa50c0b5ae553932732ec83ce1f1ec594b71dbafff59e6cd24695889
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / gdelt_regkg-1.0.2-py3-none-any.whl

Download URL gdelt_regkg-1.0.2-py3-none-any.whl
Size 369.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
df23583094c237a4e79633078f0cbbe616c97a1672a2961b7997d6c8d63b7d65
BLAKE2b-256 checksum
How to use checksums
5e3f6d8c79a0c618e91304dacd280fca92bb56ebecb7e987ad4b316b4320ec40
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.2 This release

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page