Skip to main content

localis

Fast, offline access to comprehensive data for countries, subdivisions, cities and the macroregions countries sit in. Built on ISO 3166, GeoNames and Unicode CLDR datasets (updated monthly) with support for exact lookups, filtering, and fuzzy search.

Features

  • 🌍 281 countries (31 historic) sourced and merged from ISO 3166-1, ISO 3166-3, and GeoNames
  • 🗺️ 51,711 subdivisions sourced and merged from ISO 3166-2 and GeoNames
  • 🏙️ 235,915 cities sourced from GeoNames cities500.txt
  • 🌐 34 macroregions (5 regions, 23 subregions, 6 groupings) sourced from Unicode CLDR, with every current country placed in them
  • 🔍 Typo-tolerant search: with a typo in the query, the intended record ranks first for 99.3% of countries, 85.2% of subdivisions and 91.6% of cities
  • 📌 Aliases - support for colloquial, historic and alternate names

Installation

pip install localis

Quick Start

import localis

# Countries
country = localis.countries.lookup("US")
print(country.name)  # "United States"

# Subdivisions
state = localis.subdivisions.lookup("US-CA")
print(state.name)  # "California"

# Fuzzy search
results = localis.countries.search("Austrlia")  # Typo-tolerant
print(results[0][0].name)  # "Australia"

Countries

Get

import localis

# By localis ID
country = localis.countries.get(1)

Returns: Country object or None

Lookup

# By alpha-2 code
country = localis.countries.lookup("GB")

# By alpha-3 code
country = localis.countries.lookup("GBR")

# By numeric code
country = localis.countries.lookup(826)

Returns: Country object or None

lookup() matches codes only (alpha-2, alpha-3, numeric). Common abbreviations that aren't ISO codes, such as "UK" for the United Kingdom, are found by filter(name=...) and search().

Filter

# Exact name match (searches name, official_name, common_name, and aliases)
results = localis.countries.filter(name="Canada")

# General query across all fields
results = localis.countries.filter(name="United", limit=5)

# By macroregion: a region, subregion or grouping, by name or code
results = localis.countries.filter(macroregion="Western Europe")
results = localis.countries.filter(macroregion="EU")

Returns: list[Country]

# Typo-tolerant search
results = localis.countries.search("Germny", limit=5)

for country, score in results:
    print(f"{country.name}: {score}")
# Output:
# Germany: 0.951
# Guernsey: 0.714
# ...

Returns: list[tuple[Country, float]] - sorted by similarity score

Iteration

# Iterate over all countries
for country in localis.countries:
    print(country.name)

# Get count
total = len(localis.countries)

Historic Countries

ISO 3166-3 withdrawn countries (Czechoslovakia, Serbia and Montenegro, Netherlands Antilles, and others) are included in the dataset but excluded from filter(), search(), and iteration by default.

localis.countries.include_historic   # False
len(list(localis.countries))         # 250

# Include historic entries
localis.countries.set_include_historic(True)
len(list(localis.countries))         # 281

localis.countries.set_include_historic(False)

The toggle applies to every thread using localis.countries, so set it before sharing the registry between threads (see Concurrency).

get() and lookup() always resolve historic entries regardless of the toggle. ISO reused alpha2/alpha3/numeric codes across different withdrawn countries over time (e.g. CS was both Czechoslovakia and, decades later, Serbia and Montenegro), so lookup() only resolves a historic entry by its unique alpha_4 withdrawal code, never by bare alpha2/alpha3/numeric:

Country Object

country = localis.countries.lookup("US")

country.id            # Database ID
country.name          # "United States" - ISO 3166-1 name, as published
country.official_name # "United States of America" - ISO 3166-1 official name, or None where ISO has none
country.common_name   # None - common name from Debian iso-codes where it differs (e.g. "South Korea" for "Korea, Republic of"), otherwise None
country.alpha2        # "US"
country.alpha3        # "USA"
country.geonames_id   # 6252001
country.numeric       # 840
country.aliases       # tuple[str, ...] - Alternate names
country.flag          # "🇺🇸" - Unicode flag emoji
country.historic      # HistoricInfo | None - set only for withdrawn ISO 3166-3 countries
country.macroregions  # tuple[MacroregionBase, ...] - CLDR path, region then subregion: (Americas, Northern America); () for most historic countries
country.groupings     # tuple[MacroregionBase, ...] - CLDR groupings the country belongs to: (North America, United Nations)

country = localis.countries.lookup("CSHH")  # Czechoslovakia
country.historic.alpha_4            # "CSHH"
country.historic.withdrawal_date    # "1993-01-01"
country.historic.comment            # str | None

# Utility methods
country.to_dict()     # Convert to dictionary
country.json()        # Convert to JSON string

Subdivisions

Get by ID

import localis

# By localis ID
subdivision = localis.subdivisions.get(1)

Returns: Subdivision object or None

Lookup by identifier

# By ISO code (country-subdivision)
subdivision = localis.subdivisions.lookup("US-CA")

# By GeoNames code
subdivision = localis.subdivisions.lookup("US.CA")

Returns: Subdivision object or None

Filter

# Exact name match
results = localis.subdivisions.filter(name="California")

# By subdivision type
results = localis.subdivisions.filter(type="state")

# By country
results = localis.subdivisions.filter(country="United States")

# By admin level (1 = states/provinces, 2 = counties/districts)
results = localis.subdivisions.filter(admin_level=1)

# Combine multiple filters (AND logic)
results = localis.subdivisions.filter(
    country="US",
    type="state",
    limit=10
)

Returns: list[Subdivision]

Fuzzy Search

results = localis.subdivisions.search("Californa", limit=3)

for subdivision, score in results:
    print(f"{subdivision.name}: {score}")
# California: 0.94
# Baja California: 0.8
# ...

Returns: list[tuple[Subdivision, float]]

Subdivision Object

subdivision = localis.subdivisions.lookup("US-CA")

subdivision.id              # Database ID
subdivision.name            # "California"
subdivision.geonames_code   # "US.CA"
subdivision.iso_code        # "US-CA"
subdivision.type            # "State"
subdivision.admin_level     # 1
subdivision.parent          # SubdivisionBase | None - Parent subdivision
subdivision.country         # CountryBase object
subdivision.aliases         # tuple[str, ...] - Alternate names

# Utility methods
subdivision.to_dict()       # Convert to dictionary
subdivision.json()          # Convert to JSON string

Cities

Get by ID

import localis

# By localis ID
city = localis.cities.get(1)

Returns: City object or None

Lookup by GeoNames id

# By GeoNames ID
city = localis.cities.lookup(5128581)

Returns: City object or None

Filter

# Exact name match
results = localis.cities.filter(name="Los Angeles")

# By country name or alpha2/alpha 3 code
results = localis.cities.filter(country="United States", limit=10)

# By subdivision name or ISO/GeoNames code
results = localis.cities.filter(subdivision="California", limit=10)

# Combine filters (AND logic)
results = localis.cities.filter(
    country="US",
    subdivision="California",
    limit=20
)

Returns: list[City]

Fuzzy Search

results = localis.cities.search("Los Angelos", limit=5)

for city, score in results:
    print(f"{city.name}, {city.country.name}: {score}")

Returns: list[tuple[City, float]] - sorted by similarity score

Population Threshold

# Narrow the cache and all indexes to cities with population >= 15000
localis.cities.set_population_threshold(15000)

# Check the current threshold
localis.cities.population_threshold  # 15000

# Reset back to the full dataset
localis.cities.set_population_threshold(None)

cities is fully lazy-loaded, nothing is read from disk until first access. Call set_population_threshold() before that first access (before any .get(), .lookup(), .filter(), .search(), or .force_cache() call) so the registry only ever loads the narrowed dataset. Calling it after the cache or indexes are already built still works, but it invalidates them, so the next access rebuilds the caches from scratch at the new threshold. The threshold applies to every thread using localis.cities, so set it before sharing the registry between threads (see Concurrency).

Returns: None

City Object

city = localis.cities.lookup(5128581) # GeoNames ID

city.id              # Database ID
city.geonames_id     # 5128581
city.name            # "New York"
city.subdivisions    # list[SubdivisionBase] - ordered by admin_level ascending
city.country         # CountryBase object
city.population      # 8804190
city.lat             # 40.71427
city.lng             # -74.00597

# Utility methods
city.to_dict()       # Convert to dictionary
city.json()          # Convert to JSON string

Macroregions

Get

# By localis ID
region = localis.macroregions.get(1)

Returns: Macroregion object or None

Lookup

# By code: M49 numeric codes are zero-padded strings, so lookup(9) finds nothing
oceania = localis.macroregions.lookup("009")
eu = localis.macroregions.lookup("EU")

# By name
western_europe = localis.macroregions.lookup("Western Europe")

Returns: Macroregion object or None

Iteration

for macroregion in localis.macroregions:
    print(macroregion.name)

total = len(localis.macroregions)

Macroregion Object

macroregion = localis.macroregions.lookup("155")

macroregion.id        # Database ID
macroregion.name      # "Western Europe" - CLDR English name
macroregion.code      # "155" - M49 numeric code as a string, or CLDR's letter code ("QO", "EU")
macroregion.type      # "region", "subregion" or "grouping"
macroregion.parent    # MacroregionBase | None - a subregion's region, or the region CLDR files a grouping under

# Utility methods
macroregion.to_dict() # Convert to dictionary
macroregion.json()    # Convert to JSON string

Base Objects

Basic versions of country, subdivision and macroregion when nested.

CountryBase Object

nested_country = subdivision.country

nested_country.id
nested_country.name
nested_country.alpha2
nested_country.alpha3
nested_country.geonames_id

SubdivisionBase Object

nested_sub = city.subdivisions[0]

nested_sub.id
nested_sub.name
nested_sub.geonames_code
nested_sub.iso_code
nested_sub.type
nested_sub.admin_level

MacroregionBase Object

nested_macroregion = country.macroregions[0]

nested_macroregion.id
nested_macroregion.name
nested_macroregion.code
nested_macroregion.type

Performance

Caching

All registries and their indexes are lazy-loaded on first use, incurring a cold start cost on whichever call touches them first. Any registry's dataset and indexes can be pre-loaded with .force_cache() to avoid this during queries, or you can simply access the registry/method to trigger the lazy loading upfront.

A registry's dataset also loads the datasets it references, if they aren't cached yet. Countries load macroregions, subdivisions load countries, and cities load subdivisions and countries. Only those datasets load, not their indexes. The subdivisions and cities tables below exclude them, so a cold first call on cities also pays for the subdivisions and countries datasets.

Countries (281)

Component Load Time Memory
Dataset ~1ms 215KB
Lookup index < 1ms 41KB
Filter index < 1ms 119KB
Search index ~4ms 381KB
Combined ~6ms 756KB

Subdivisions (51,711)

Component Load Time Memory
Dataset ~85ms 16.4MB
Lookup index ~16ms 4.3MB
Filter index ~120ms 13.8MB
Search index ~62ms 12.8MB
Combined ~284ms 47.4MB

Cities (235,915)

⚠️ Memory-intensive. Fully caching cities and its indexes adds 114.4MB of memory. Calling localis.cities.force_cache() loads all of it upfront. You can call cities.set_population_threshold(n) before first access as a lever to control the memory footprint.

Component Load Time Memory
Dataset ~341ms 28.1MB
Lookup index ~74ms 1.8MB
Filter index ~737ms 48.5MB
Search index ~150ms 36.0MB
Combined ~1.30s 114.4MB

At a threshold of 15,000, cities drops from 235,915 to 34,171 and memory drops from 114.4MB to 26.5MB.

Full Cache: ~1.58s load time, 162.5MB memory for all datasets and indexes

Benchmarks

Per-call query latency on warm caches, median (95th percentile), and fuzzy search accuracy on mangled/misspelled queries: how often the right entry appears in the top 10 results, and how often it's the top result.

Registry Get Lookup Filter Search (p50/p95) Search Accuracy (top 10) Top Result
Countries 0.0091ms (0.0125ms) 0.0126ms (0.0166ms) 0.018ms (0.0237ms) 2.12ms (4.92ms) 100.0% 99.3%
Subdivisions 0.0085ms (0.0105ms) 0.0136ms (0.0168ms) 0.0185ms (0.0363ms) 3.33ms (5.77ms) 97.0% 85.2%
Cities 0.0142ms (0.0181ms) 0.012ms (0.0146ms) 0.0238ms (0.0852ms) 6.82ms (11.8ms) 98.7% 91.6%

Accuracy tested on 5,000 mangled-query samples per registry; cities' search additionally includes city + admin1 context. Load times, memory and latency are generated by tests/analysis/footprint.py and tests/analysis/benchmarks.py, last measured on 11th Gen Intel(R) Core(TM) i7-1165G7 @ 2.80GHz with Python 3.14.7.


Data Sources

Data in this project is kept current monthly from the following sources:

  • Countries
  • Subdivisions
  • Cities
  • Macroregions
    • Unicode CLDR territory containment and English territory names

docs/methodology.md is a complete, falsifiable account of how each dataset is built: the rules that combine these sources, how the results were validated, and where they are known to be wrong. unmerged_subdivisions.md lists every ISO subdivision currently without a GeoNames counterpart, regenerated on every ingest run.

Concurrency

Registries are safe to share across threads. get(), lookup(), filter(), search() and iteration only read shared data, and the first access that loads a dataset or index does so under the registry's lock, so threads reaching a cold registry together load it once.

Two settings change shared state for every thread: cities.set_population_threshold() and countries.set_include_historic(). Configure them before the registry is shared between threads, never while other threads are querying it.

Batch searching

localis doesn't parallelize batches for you, since the right approach depends on your Python build, memory budget and surrounding executor. On free-threaded Python (3.14t), a thread pool searches in parallel:

from concurrent.futures import ThreadPoolExecutor
import localis

localis.cities.force_cache()  # load once, before the threads start
with ThreadPoolExecutor() as pool:
    results = list(pool.map(localis.cities.search, queries))

On a standard Python build the same code is correct but runs one search at a time, because the GIL lets only one thread run Python code at once. To search in parallel there, use processes. Each worker loads its own copy of the data (up to 114.4MB for cities), so apply any settings in the worker's initializer, and search through a module-level function, since a registry itself can't be sent to a process:

from concurrent.futures import ProcessPoolExecutor
import localis

def init_worker():
    localis.cities.set_population_threshold(15000)  # repeat any setting the parent uses

def search_city(query):
    return localis.cities.search(query)

if __name__ == "__main__":  # workers import this module, so the pool only starts in the parent
    with ProcessPoolExecutor(initializer=init_worker) as pool:
        results = list(pool.map(search_city, queries, chunksize=500))

Data licensing

The shipped data is derived from these sources and remains subject to their licenses: ISO 3166 data via iso-codes (LGPL-2.1-or-later), GeoNames (CC BY 4.0), Wikidata (CC0), and Unicode CLDR (Unicode License v3).


Requirements

  • Python 3.11+
  • rapidfuzz - Fast fuzzy string matching
  • unidecode - Unicode text normalization

License

Code: MIT. Data: subject to its sources' licenses, listed under Data licensing.


Why localis

localis began with some database cleanup. I found myself writing mountains of bespoke code to parse inconsistent, dirty data with pycountry, GeoNames and Wikidata to name a few (Google Places was not in the budget). When the pipeline was complete and the data cleaned, I realized this mountain of code could be useful for others who might need a reliable offline geo-data solution, so here we are! I hope you find it useful and please don't hesitate to contribute or report any issues.


Contributing

Pull requests welcome Report issues

Support this project:

Metadata

Release files for localis 2.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for localis 2.1.0
File Size Uploaded
localis-2.1.0.tar.gz 21.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for localis 2.1.0
File Interpreter ABI Platform
localis-2.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 42.5 MB

Release files / localis-2.1.0.tar.gz

Download URL localis-2.1.0.tar.gz
Size 21.2 MB
Tags Source
SHA-256 checksum
How to use checksums
7f9f2080cc6820399a3c79a742e0ed37ae873cdc79a36a9972afc381f65b2f05
BLAKE2b-256 checksum
How to use checksums
61f8709f2523572f88bd040bcc732f8cdc92ed5dc88c8161d397f4418cac3491
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / localis-2.1.0-py3-none-any.whl

Download URL localis-2.1.0-py3-none-any.whl
Size 21.4 MB
Tags Python 3
SHA-256 checksum
How to use checksums
bcb629be25d5fa8f86382bef720258eb2c1a535ea03cdc787d7bd2827e45de79
BLAKE2b-256 checksum
How to use checksums
3088262a49c04f95f755f7f1a3c7eaf69687e3d929e6da69cff858a4190f59b2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page