owid-catalog
A Pythonic library for working with OWID data.
The owid-catalog library is the foundation of Our World in Data's data management system. It provides:
- Data APIs: Access OWID's published data through unified client interfaces
- Data Structures: Enhanced pandas DataFrames with rich metadata support
Installation
pip install owid-catalog
Quick Examples
Accessing OWID Data
from owid.catalog import fetch, search
# Search for charts (default)
charts = search("population")
tb = charts[0].fetch()
# Fetch data from OWID Chart at ourworldindata.org/grapher/life-expectancy
tb = fetch("life-expectancy")
# Search for tables
tables = search("population", kind="table", namespace="un")
tb = tables[0].fetch()
# Search indicators (using semantic search)
search("renewable energy", kind="indicator")
Working with Data Structures
from owid.catalog import Table
from owid.catalog import processing as pr
# Tables are pandas DataFrames with metadata
tb = Table(df, metadata={"short_name": "population"})
# Metadata propagates through operations
tb_filtered = tb[tb["year"] > 2000] # Keeps metadata
tb_merged = pr.merge(tb1, tb2, on="country") # Merges metadata
Documentation
For detailed documentation, see:
- API Reference: ChartsAPI, IndicatorsAPI, TablesAPI
- Data Structures: Dataset, Table, Variable, metadata handling
- Full Documentation: Complete library documentation
- Agent skill: guide for AI agents using the library for data analysis
Architecture
graph TB
etl -->|reads| snapshot[upstream datasets]
etl -->|generates| s3[data catalog]
catalog[owid-catalog] -->|queries| s3
This library is part of OWID's ETL project, which contains recipes for all datasets we publish.
Development
You need Python 3.11+, uv and make installed. Clone the repo, then you can simply run:
# run all unit tests and CI checks
make test
# watch for changes, then run all checks
make watch
Maintainer notes — how the version is bumped, how a release reaches PyPI, and which checks actually cover this directory — are in DEVELOPMENT.md.
Changelog
v1.2.6
- Documentation rendered from metadata (new
core.docsmodule)Dataset.readme()renders a Markdown README for a dataset: its description, one block per table (citation, indicators, sources), its changelog, the processing note, license and download links- New
Table.sourcestable of the distinct origins behind a table's columns;Table.codebookis rebuilt on the same helpers, and itssourcecolumn uses the same labels
- Dataset changelog: new
DatasetMeta.changelog, a list ofChangelogEntry(date, changes), set fromdataset.changelogin.meta.yml. Dates are stored asYYYY-MM-DDstrings, and a malformed entry raises where it is set. The README renders it as its own "Changelog" section, and it never enters the description - Schema.org JSON-LD
- Descriptions over Google Dataset Search's 5,000-character limit raise instead of producing a record Google rejects
- Each table's nested Dataset repeats the
creator, since Google validates nested Datasets on their own dataset_keywords()is public (was_keywords)
- Docstrings say "data producer" instead of "data provider"
v1.2.5
- Remove the leftover
multidimchannel fromCHANNELincore.datasetsandcore.paths. MDIM configs were written as datasets underdata/multidim/for a short while in 2025; they have lived outside the data catalog since, so the channel only made the publish job look for an empty folder and write an empty index file on every deploy - Render metadata Jinja in a
SandboxedEnvironment, so a producer-supplied value containing<<or<%cannot walk__class__/__subclasses__on whatever runs the ETL - Fail loudly on a character-exploded
description_key(a markdown string that was iterated character by character):Markdownis now astrsubclass that raisesTypeErroron iteration, alongside a newvalidate_description_key_list()sanity check - New
s3_utils.object_exists()to check for an object without downloading it
v1.2.4
- Python support
- Add Python 3.14, drop Python 3.10 (
requires-python = ">=3.11, <3.15") - Drop the
typing_extensionsfallbacks forSelf,RequiredandNotRequired
- Add Python 3.14, drop Python 3.10 (
description_keybecomes free-form markdownVariableMeta.description_keyis now a markdown string- A list of bullet points is still accepted (items may carry per-item Jinja) and is converted to a markdown list after rendering — the grapher only ever sees a string
- New
description_key_to_string()inowid.catalog.core.meta, reproducing how the grapher rendered those lists before
- Display metadata
- Add
display.timeInterval(day,week,month,quarter,year,decade) - Remove the deprecated
display.yearIsDay
- Add
v1.2.3
combine_indicators_processing_leveltolerates unrendered Jinja templates inprocessing_levelinstead of asserting on an unknown level — when a template is combined with a literal it overstates rather than understates the result, sinceprocessing_levelfeeds licensing downstreams3_utils.upload()acceptscontent_typeandcache_control- Catalog JSON-LD (
schema_org): license and keyword fixes,citationdropped - Dependency bumps for security advisories
v1.2.2
- Add
ownerstoDatasetMeta - New
Dataset.update_metadata_from_dict()andyaml_metadata.update_metadata_from_dict(), for callers that already hold parsed metadata rather than a YAML path (used by the Owl runner) - New
schema_orgmodule emitting schema.org JSON-LD for catalog datasets and tables, with follow-ups: stable short landing-page URLs, table descriptions filled from existing metadata, an explicit dataset description requirement, and unrendered Jinja templates kept out of the output - YAML variable checks accept a long-format base name as a match for pivoted
{base}__{dim}_{value}columns, so along_to_wideoverride block is no longer flagged as a typo - Republished to PyPI after the
tyupgrade, with no functional change of its own
v1.2.1
- Send
User-Agent: owid-catalog/<version> (python <x.y.z>)on every outbound HTTP call, via a sharedrequests.Sessioninowid.catalog.api.utils(plusSTORAGE_OPTIONSfor the pandas reads that cannot take a session) prune_dict()keeps explicitly-empty values for keys listed inKEEP_IF_EMPTY:chartTypes: []is a meaningful "render no chart-type toggles" override, not the same as an absent key
v1.2.0
- Remove legacy
Sourcemetadata (origins only)- Removed
Sourceclass fromowid.catalog.core.meta - Removed
sourcesfield fromVariableMetaandDatasetMeta(useoriginsinstead) - Removed
if_source_existsparameter fromDataset.update_metadata(useif_origins_exist) - Removed
get_unique_sources_from_indicatorshelper - Removed
sourcesaggregation fromcombine_indicators_metadata
- Removed
v1.1.0
- Remove processing log feature
- Removed
ProcessingLogandprocessing_logmodule fromowid.catalog.core - Removed
combine_indicators_processing_logshelper - Removed
update_log/amend_logmethods onIndicator - Removed processing-log tracking from
Tablearithmetic operations (__add__,__sub__,__mul__, etc.) - Removed
processing_logfield fromVariableMeta
- Removed
v1.0.1
- ResponseSet ergonomics
- Remove deprecated
ResponseSet.resultsproperty (use.itemsinstead) - Add
.to_dict()method for serializing results to plain dicts (useful for AI/LLM context windows) - Add
all_fieldsparameter to.to_frame()to temporarily override display mode without mutating instance state
- Remove deprecated
v1.0.0
- New unified Client API
owid.catalog.Clientas single entry point withChartsAPI,IndicatorsAPI,TablesAPI- Quick access via
search()andfetch()convenience functions - Rich result types:
ChartResult,IndicatorResult,TableResultwithResponseSetcontainer
- Charts API
- Fetch chart data by slug, URL, or slug with query params
- Parse chart slugs from grapher/explorer URLs via
parse_chart_slug() - Explorer best-effort fetching with graceful error handling
set_ui_advanced()/set_ui_basic()for display configuration
- Tables API
- Search catalog by table, namespace, version, dataset, and channel
- Fetch tables directly by catalog path
- Embedded catalog index with local caching
- Indicators API
- Semantic search via
search.owid.iovector embeddings - Sort by relevance (similarity + popularity blend) or similarity only
fetch()for single-column indicator orfetch_table()for the full table
- Semantic search via
- Search & discovery
- Fuzzy, exact, contains, and regex matching modes
.latest()filtering to keep only newest versions- Popularity scores (0.0-1.0) from analytics views, results sorted by popularity
refresh_indexparameter to force catalog index reload
- Data structures integration
- All
fetch()methods returnowid.catalog.Tablewith full metadata CatalogPathhelper for parsing catalog paths- Lazy loading with
load_data=Falsefor deferred data access
- All
- Library reorganization
- Restructured into
owid.catalog.core(data structures) andowid.catalog.api(remote access) catalog.find()deprecated in favor ofClient().tables.search()(backwards compat maintained)- Legacy code moved to
owid.catalog.api.legacy - New dependencies:
pydanticv2.0+
- Restructured into
- Private data support
- Private datasets served from separate R2 bucket
- API can fetch private data from private bucket
- Performance
- Vectorized operations replacing
iterrows()in TablesAPI - Embedded catalog index loading (removed ETLCatalog dependency)
- Modularized search into helper methods
- Vectorized operations replacing
- Other
- Thumbnail display in
ResponseSetfor chart results - JSON output format support
- Comprehensive exception handling:
ChartNotFoundError,LicenseError - API URLs immutable with Pydantic
Field(frozen=True)
- Thumbnail display in
See previous versions
v0.4.5
- Allow both
tableanddatasetparameters infind()(they can now be used together) - Migrate from pyright to ty type checker for improved type checking
v0.4.4
- Enhanced
find()with better search capabilities:- Case-insensitive search by default (use
case=Truefor case-sensitive) - Regex support enabled by default for
tableanddatasetparameters - New fuzzy search with
fuzzy=True- typo-tolerant matching sorted by relevance - Configurable fuzzy threshold (0-100) to control match strictness
- Case-insensitive search by default (use
- New dependency:
rapidfuzzfor fuzzy string matching
v0.4.3
- Fixed minor bugs
v0.4.0
- Highlights
- Support for Python 3.10-3.13 (was 3.11-3.13)
- Drop support for Python 3.9 (breaking change)
- Others
- Deprecate Walden.
- Dependencies: Change
rdataforpyreadr. - Support: indicator dimensions.
- Support: MDIMs.
- Switched from Poetry to UV package manager.
- New decorator
@keep_metadatato propagate metadata in pandas functions.
- Fixes:
Table.apply,groupby.apply, metadata propagation, type hinting, etc.
v0.3.11
- Add support for Python 3.12 in
pypackage.toml
v0.3.10
- Add experimental chart data API in
owid.catalog.charts
v0.3.9
- Switch from isort & black & fake8 to ruff
v0.3.8
- Pin dataclasses-json==0.5.8 to fix error with python3.9
v0.3.7
- Fix bugs.
- Improve metadata propagation.
- Improve metadata YAML file handling, to have common definitions.
- Remove
DatasetMeta.origins.
v0.3.6
- Fixed tons of bugs
processing.pymodule with pandas-like functions that propagate metadata- Support for Dynamic YAML files
- Support for R2 alongside S3
v0.3.5
- Remove
catalog.frames; useowid-repackpackage instead - Relax dependency constraints
- Add optional
channelargument toDatasetMeta - Stop supporting metadata in Parquet format, load JSON sidecar instead
- Fix errors when creating new Table columns
v0.3.4
- Bump
pyarrowdependency to enable Python 3.11 support
v0.3.3
- Add more arguments to
Table.__init__that are often used in ETL - Add
Dataset.update_metadatafunction for updating metadata from YAML file - Python 3.11 support via update of
pyarrowdependency
v0.3.2
- Fix a bug in
Catalog.__getitem__() - Replace
mypytype checker bypyright
v0.3.1
- Sort imports with
isort - Change black line length to 120
- Add
grapherchannel - Support path-based indexing into catalogs
v0.3.0
- Update
OWID_CATALOG_VERSIONto 3 - Support multiple formats per table
- Support reading and writing
parquetfiles with embedded metadata - Optional
repackargument when adding tables to dataset - Underscore
| - Get
versionfield fromDatasetMetainit - Resolve collisions of
underscore_tablefunction - Convert
versiontostrand load jsondimensions
v0.2.9
- Allow multiple channels in
catalog.findfunction
v0.2.8
- Update
OWID_CATALOG_VERSIONto 2
v0.2.7
- Split datasets into channels (
garden,meadow,open_numbers, ...) and make garden default one - Add
.find_latestmethod to Catalog
v0.2.6
- Add flag
is_publicfor public/private datasets - Enforce snake_case for table, dataset and variable short names
- Add fields
published_byandpublished_atto Source- Added a list of supported and unsupported operations on columns
- Updated
pyarrow
v0.2.5
- Fix ability to load remote CSV tables
v0.2.4
- Update the default catalog URL to use a CDN
v0.2.3
- Fix methods for finding and loading data from a
LocalCatalog
v0.2.2
- Repack frames to compact dtypes on
Table.to_feather()
v0.2.1
- Fix key typo used in version check
v0.2.0
- Copy dataset metadata into tables, to make tables more traceable
- Add API versioning, and a requirement to update if your version of this library is too old
v0.1.1
- Add support for Python 3.8
v0.1.0
- Initial release, including searching and fetching data from a remote catalog
Metadata
Release files for owid-catalog 1.2.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| owid_catalog-1.2.6.tar.gz | 356.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| owid_catalog-1.2.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 502.4 kB
Release files / owid_catalog-1.2.6.tar.gz
| Download URL | owid_catalog-1.2.6.tar.gz |
|---|---|
| Size | 356.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e1b6a8feb2208d641f701462c85f1acbd9b1b96a2face2aa52cbd25d39745c7e
|
|
BLAKE2b-256 checksum How to use checksums |
258db4caf7e9e60c4adb7dfdf22564fb1e8d56b3932ac1522c2d23a82e17575a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / owid_catalog-1.2.6-py3-none-any.whl
| Download URL | owid_catalog-1.2.6-py3-none-any.whl |
|---|---|
| Size | 145.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d898f7ff34d391596899bbb1188f3d3f573812dcfa7a87016033a1570a17f693
|
|
BLAKE2b-256 checksum How to use checksums |
20635ec5cafdd88840e47e2dc46a495afc8b02146df687ea72951dd8b9033d9e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|