gdelt-regkg
Package for reconstructing GDELT GKG files from article text and metadata.
gdelt-regkg converts prepared article bodies and supplied metadata into GKG records. It can be used for reconstructing gaps in GKG coverage, processing your own news datasets, or extract individual fields for analysis. To make it work you just need article text and some metadata, while the library extracts and formats the fields automatically. You can generate full GKG entry or extract data only for specific fields.
The row structure, metadata formatting and field encodings are documented in GDELT Global Knowledge Graph Codebook. Extraction from text approximates GDELT's processing, so themes, counts, entities, tone and other inferred values can differ from native GKG results.
See the field documentation and measured agreement, resource and benchmark tools, and changelog.
Installation
Requires Python 3.12 or newer. Install the library and its English model, then prepare the extraction dictionaries:
python -m pip install gdelt-regkg
python -m spacy download en_core_web_sm
gdelt-regkg-build-resources --sources-dir downloads/resources --download-sources --output local-resources
Installation and setup explains what is downloaded, how to run your first extraction, and how to add larger models, GCAM dictionaries or translation.
Usage
Extract individual fields
Each field has a standalone function, so you do not need to generate a full GKG row.
from gdelt_regkg.fields import extract_counts, extract_themes, extract_tone
text = "Flooding killed 12 people and displaced 300 residents."
counts = extract_counts(text)
# CRISISLEX_T03_DEAD#12#people#0######;KILL#12#people#0######;AFFECT#300#residents#0######;CRISISLEX_CRISISLEXREC#300#residents#0######;CRISISLEX_T09_DISPLACEDRELOCATEDEVACUATED#300#residents#0######;DISPLACED#300#residents#0######;
themes = extract_themes(text)
# CRISISLEX_C06_WATER_SANITATION;NATURAL_DISASTER_FLOODING;CRISISLEX_T03_DEAD;KILL;AFFECT;CRISISLEX_CRISISLEXREC;CRISISLEX_T09_DISPLACEDRELOCATEDEVACUATED;DISPLACED;MANMADE_DISASTER_IMPLIED;
tone = extract_tone(text)
# -12.5,0,12.5,12.5,12.5,0,8
Functions return GKG-formatted cell strings and work with pandas, for example df["themes"] = df["text"].map(extract_themes). pandas package is optional and must be installed separately.
Tone, themes, geographic references and GCAM require locally prepared profiles. The default directory is local-resources/; select another with resources_dir= in Python or --resources-dir in the CLI. Explicit resource objects override directory defaults. Counts and name rules are available directly from the package.
Generate and save a row
from gdelt_regkg import SpacyNER, generate_gkg, write_gkg
ner = SpacyNER.from_model("en_core_web_sm", name_policy="surface") # Load once and reuse.
record = generate_gkg(
"Volunteers rescued 12 people in Paris, France.",
identifier="https://example.org/story",
published_at="20261001090023", # Known precise publication time, stored in extra XML.
batch_time="20261001091500",
sequence=1,
ner=ner,
resources_dir="local-resources",
sharing_image="",
related_images=[],
social_image_embeds=[],
social_video_embeds=[],
extras={},
strict=True,
)
with open("example.gkg.csv", "w", encoding="utf-8", newline="") as stream:
write_gkg([record], stream)
for column in ("V1COUNTS", "V1LOCATIONS", "V2.1AMOUNTS"):
print(f"{column}: {record[column]}")
V1COUNTS: CRISISLEX_T08_MISSINGFOUNDTRAPPEDPEOPLE#12#people#4#Paris, France (General), France#FR#FR00#48.8667#2.33333#-1456928;
V1LOCATIONS: 4#Paris, France (General), France#FR#FR00#48.8667#2.33333#-1456928;1#France#FR#FR#46#2#FR
V2.1AMOUNTS: 12,people,19
These are three cells from the generated record. example.gkg.csv contains the complete 27-column GKG entry, including themes, tone, GCAM and enhanced fields. The generator resolves locations from NER and adds the nearest resolved location to each count.
The generator returns a dictionary of 27 fields. This example explicitly supplies a recognizer and empty media/extras metadata, so strict=True verifies a complete row. Missing supplied metadata or model-dependent fields without ner remain None and produce a warning. allow_incomplete=True exports those unavailable values as blank cells. Output is headerless and tab-delimited, even when the filename ends in .csv (GDELT specifics).
For non-English articles and the full GCAM profile, use the complete pipeline example in the setup guide.
Field implementation map
Columns appear in output order. Model-dependent fields require an installed recognizer; missing supplied metadata remains None.
| Column | Standalone function | Status |
|---|---|---|
| GKGRECORDID | format_record_id |
Supplied metadata |
| V2.1DATE | format_date |
Batch timestamp |
| V2SOURCECOLLECTIONIDENTIFIER | format_source_collection |
Supplied metadata |
| V2SOURCECOMMONNAME | source_common_name |
Explicit label or hostname |
| V2DOCUMENTIDENTIFIER | document_identifier |
Supplied metadata |
| V1COUNTS | extract_counts |
Experimental approximation |
| V2.1COUNTS | extract_enhanced_counts |
Same counts with number offsets |
| V1THEMES | extract_themes |
Rule-based approximation |
| V2ENHANCEDTHEMES | extract_enhanced_themes |
Same theme mentions with phrase offsets |
| V1LOCATIONS | extract_locations |
Offline reference gazetteer and NER approximation |
| V2ENHANCEDLOCATIONS | extract_enhanced_locations |
Same resolved places with ADM2 and offsets |
| V1PERSONS | extract_persons |
spaCy NER approximation |
| V2ENHANCEDPERSONS | extract_enhanced_persons |
Same person mentions with offsets |
| V1ORGANIZATIONS | extract_organizations |
spaCy NER approximation |
| V2ENHANCEDORGANIZATIONS | extract_enhanced_organizations |
Same organization mentions with offsets |
| V1.5TONE | extract_tone |
Benchmark-refined dictionary approximation |
| V2.1ENHANCEDDATES | extract_enhanced_dates |
Explicit English dates with resolutions 1–4 |
| V2GCAM | extract_gcam |
Dictionary category counts and weighted scores; coverage depends on local dictionaries |
| V2.1SHARINGIMAGE | format_sharing_image |
Supplied image URL |
| V2.1RELATEDIMAGES | format_related_images |
Supplied image URLs |
| V2.1SOCIALIMAGEEMBEDS | format_social_image_embeds |
Supplied social-post URLs |
| V2.1SOCIALVIDEOEMBEDS | format_social_video_embeds |
Supplied video URLs |
| V2.1QUOTATIONS | extract_quotations |
Quoted spans and reporting-verb approximation |
| V2.1ALLNAMES | extract_all_names |
Broad named entities and proper-noun spans |
| V2.1AMOUNTS | extract_amounts |
Numeric/textual quantities with bounded objects |
| V2.1TRANSLATIONINFO | format_translation_info |
Supplied metadata |
| V2EXTRASXML | extract_extras_xml |
Custom extension data, including supplied metadata |
See the field documentation for implementation status, research, and usage for each field.
Contributing
Use the GitHub issue forms to report bugs, extraction differences or feature requests. Contributing describes local checks, pull requests and evaluation on data separate from rule development.
Security
Report vulnerabilities through the repository's Security → Report a vulnerability form rather than a public issue. Include the affected version, a minimal reproduction and impact. Do not include credentials or article bodies you cannot share. If private reporting is unavailable, use a private contact listed on the maintainer's GitHub profile.
Security fixes target the latest release. Ordinary extraction disagreements and model errors belong in the issue forms.
Disclaimer
This is an independent implementation, unaffiliated with and not endorsed by the GDELT Project. It produces approximate GKG fields, not guaranteed copies of GDELT's outputs. Results depend on article text, translation, models and dictionaries; validate them for your intended analysis. Differences from native GKG can also reflect different input bodies.
The code and original rules use the MIT license and are provided without warranty. MIT does not grant rights to external dictionaries, models, article bodies or generated profiles derived from them; their source terms still apply.
Metadata
Release files for gdelt-regkg 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gdelt_regkg-1.0.0.tar.gz | 438.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gdelt_regkg-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 807.6 kB
Release files / gdelt_regkg-1.0.0.tar.gz
| Download URL | gdelt_regkg-1.0.0.tar.gz |
|---|---|
| Size | 438.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
426fa1955bbcec4ab95072fd5aba746b535969f18d69f24507437750c7ec5c2f
|
|
BLAKE2b-256 checksum How to use checksums |
80e12199f315812ddf9355d4ed6bb54f586848370ec0c0cf7c1b8f52449f78f6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / gdelt_regkg-1.0.0-py3-none-any.whl
| Download URL | gdelt_regkg-1.0.0-py3-none-any.whl |
|---|---|
| Size | 369.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
357563bcd93e25b45490bf1ddc34bc43f751e97ce85c758b2c7ee1c6c8dae7c4
|
|
BLAKE2b-256 checksum How to use checksums |
a2d82081258b6e9ec917bb110d1e9c6665b077158690129c3f1c270f2be417bb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log