localis
Fast, offline access to comprehensive data for countries, subdivisions, cities and the macroregions countries sit in. Built on ISO 3166, GeoNames and Unicode CLDR datasets (updated monthly) with support for exact lookups, filtering, and fuzzy search.
Features
- 🌍 281 countries (31 historic) sourced and merged from ISO 3166-1, ISO 3166-3, and GeoNames
- 🗺️ 51,711 subdivisions sourced and merged from ISO 3166-2 and GeoNames
- 🏙️ 235,915 cities sourced from GeoNames cities500.txt
- 🌐 34 macroregions (5 regions, 23 subregions, 6 groupings) sourced from Unicode CLDR, with every current country placed in them
- 🔍 Typo-tolerant search: with a typo in the query, the intended record ranks first for 99.3% of countries, 85.2% of subdivisions and 91.6% of cities
- 📌 Aliases - support for colloquial, historic and alternate names
Installation
pip install localis
Quick Start
import localis
# Countries
country = localis.countries.lookup("US")
print(country.name) # "United States"
# Subdivisions
state = localis.subdivisions.lookup("US-CA")
print(state.name) # "California"
# Fuzzy search
results = localis.countries.search("Austrlia") # Typo-tolerant
print(results[0][0].name) # "Australia"
Countries
Get
import localis
# By localis ID
country = localis.countries.get(1)
Returns: Country object or None
Lookup
# By alpha-2 code
country = localis.countries.lookup("GB")
# By alpha-3 code
country = localis.countries.lookup("GBR")
# By numeric code
country = localis.countries.lookup(826)
Returns: Country object or None
lookup() matches codes only (alpha-2, alpha-3, numeric). Common abbreviations that aren't ISO codes, such as "UK" for the United Kingdom, are found by filter(name=...) and search().
Filter
# Exact name match (searches name, official_name, common_name, and aliases)
results = localis.countries.filter(name="Canada")
# General query across all fields
results = localis.countries.filter(name="United", limit=5)
# By macroregion: a region, subregion or grouping, by name or code
results = localis.countries.filter(macroregion="Western Europe")
results = localis.countries.filter(macroregion="EU")
Returns: list[Country]
Fuzzy Search
# Typo-tolerant search
results = localis.countries.search("Germny", limit=5)
for country, score in results:
print(f"{country.name}: {score}")
# Output:
# Germany: 0.951
# Guernsey: 0.714
# ...
Returns: list[tuple[Country, float]] - sorted by similarity score
Iteration
# Iterate over all countries
for country in localis.countries:
print(country.name)
# Get count
total = len(localis.countries)
Historic Countries
ISO 3166-3 withdrawn countries (Czechoslovakia, Serbia and Montenegro, Netherlands Antilles, and others) are included in the dataset but excluded from filter(), search(), and iteration by default.
localis.countries.include_historic # False
len(list(localis.countries)) # 250
# Include historic entries
localis.countries.set_include_historic(True)
len(list(localis.countries)) # 281
localis.countries.set_include_historic(False)
The toggle applies to every thread using localis.countries, so set it before sharing the registry between threads (see Concurrency).
get() and lookup() always resolve historic entries regardless of the toggle. ISO reused alpha2/alpha3/numeric codes across different withdrawn countries over time (e.g. CS was both Czechoslovakia and, decades later, Serbia and Montenegro), so lookup() only resolves a historic entry by its unique alpha_4 withdrawal code, never by bare alpha2/alpha3/numeric:
Country Object
country = localis.countries.lookup("US")
country.id # Database ID
country.name # "United States" - ISO 3166-1 name, as published
country.official_name # "United States of America" - ISO 3166-1 official name, or None where ISO has none
country.common_name # None - common name from Debian iso-codes where it differs (e.g. "South Korea" for "Korea, Republic of"), otherwise None
country.alpha2 # "US"
country.alpha3 # "USA"
country.geonames_id # 6252001
country.numeric # 840
country.aliases # tuple[str, ...] - Alternate names
country.flag # "🇺🇸" - Unicode flag emoji
country.historic # HistoricInfo | None - set only for withdrawn ISO 3166-3 countries
country.macroregions # tuple[MacroregionBase, ...] - CLDR path, region then subregion: (Americas, Northern America); () for most historic countries
country.groupings # tuple[MacroregionBase, ...] - CLDR groupings the country belongs to: (North America, United Nations)
country = localis.countries.lookup("CSHH") # Czechoslovakia
country.historic.alpha_4 # "CSHH"
country.historic.withdrawal_date # "1993-01-01"
country.historic.comment # str | None
# Utility methods
country.to_dict() # Convert to dictionary
country.json() # Convert to JSON string
Subdivisions
Get by ID
import localis
# By localis ID
subdivision = localis.subdivisions.get(1)
Returns: Subdivision object or None
Lookup by identifier
# By ISO code (country-subdivision)
subdivision = localis.subdivisions.lookup("US-CA")
# By GeoNames code
subdivision = localis.subdivisions.lookup("US.CA")
Returns: Subdivision object or None
Filter
# Exact name match
results = localis.subdivisions.filter(name="California")
# By subdivision type
results = localis.subdivisions.filter(type="state")
# By country
results = localis.subdivisions.filter(country="United States")
# By admin level (1 = states/provinces, 2 = counties/districts)
results = localis.subdivisions.filter(admin_level=1)
# Combine multiple filters (AND logic)
results = localis.subdivisions.filter(
country="US",
type="state",
limit=10
)
Returns: list[Subdivision]
Fuzzy Search
results = localis.subdivisions.search("Californa", limit=3)
for subdivision, score in results:
print(f"{subdivision.name}: {score}")
# California: 0.94
# Baja California: 0.8
# ...
Returns: list[tuple[Subdivision, float]]
Subdivision Object
subdivision = localis.subdivisions.lookup("US-CA")
subdivision.id # Database ID
subdivision.name # "California"
subdivision.geonames_code # "US.CA"
subdivision.iso_code # "US-CA"
subdivision.type # "State"
subdivision.admin_level # 1
subdivision.parent # SubdivisionBase | None - Parent subdivision
subdivision.country # CountryBase object
subdivision.aliases # tuple[str, ...] - Alternate names
# Utility methods
subdivision.to_dict() # Convert to dictionary
subdivision.json() # Convert to JSON string
Cities
Get by ID
import localis
# By localis ID
city = localis.cities.get(1)
Returns: City object or None
Lookup by GeoNames id
# By GeoNames ID
city = localis.cities.lookup(5128581)
Returns: City object or None
Filter
# Exact name match
results = localis.cities.filter(name="Los Angeles")
# By country name or alpha2/alpha 3 code
results = localis.cities.filter(country="United States", limit=10)
# By subdivision name or ISO/GeoNames code
results = localis.cities.filter(subdivision="California", limit=10)
# Combine filters (AND logic)
results = localis.cities.filter(
country="US",
subdivision="California",
limit=20
)
Returns: list[City]
Fuzzy Search
results = localis.cities.search("Los Angelos", limit=5)
for city, score in results:
print(f"{city.name}, {city.country.name}: {score}")
Returns: list[tuple[City, float]] - sorted by similarity score
Population Threshold
# Narrow the cache and all indexes to cities with population >= 15000
localis.cities.set_population_threshold(15000)
# Check the current threshold
localis.cities.population_threshold # 15000
# Reset back to the full dataset
localis.cities.set_population_threshold(None)
cities is fully lazy-loaded, nothing is read from disk until first access. Call set_population_threshold() before that first access (before any .get(), .lookup(), .filter(), .search(), or .force_cache() call) so the registry only ever loads the narrowed dataset. Calling it after the cache or indexes are already built still works, but it invalidates them, so the next access rebuilds the caches from scratch at the new threshold. The threshold applies to every thread using localis.cities, so set it before sharing the registry between threads (see Concurrency).
Returns: None
City Object
city = localis.cities.lookup(5128581) # GeoNames ID
city.id # Database ID
city.geonames_id # 5128581
city.name # "New York"
city.subdivisions # list[SubdivisionBase] - ordered by admin_level ascending
city.country # CountryBase object
city.population # 8804190
city.lat # 40.71427
city.lng # -74.00597
# Utility methods
city.to_dict() # Convert to dictionary
city.json() # Convert to JSON string
Macroregions
Get
# By localis ID
region = localis.macroregions.get(1)
Returns: Macroregion object or None
Lookup
# By code: M49 numeric codes are zero-padded strings, so lookup(9) finds nothing
oceania = localis.macroregions.lookup("009")
eu = localis.macroregions.lookup("EU")
# By name
western_europe = localis.macroregions.lookup("Western Europe")
Returns: Macroregion object or None
Iteration
for macroregion in localis.macroregions:
print(macroregion.name)
total = len(localis.macroregions)
Macroregion Object
macroregion = localis.macroregions.lookup("155")
macroregion.id # Database ID
macroregion.name # "Western Europe" - CLDR English name
macroregion.code # "155" - M49 numeric code as a string, or CLDR's letter code ("QO", "EU")
macroregion.type # "region", "subregion" or "grouping"
macroregion.parent # MacroregionBase | None - a subregion's region, or the region CLDR files a grouping under
# Utility methods
macroregion.to_dict() # Convert to dictionary
macroregion.json() # Convert to JSON string
Base Objects
Basic versions of country, subdivision and macroregion when nested.
CountryBase Object
nested_country = subdivision.country
nested_country.id
nested_country.name
nested_country.alpha2
nested_country.alpha3
nested_country.geonames_id
SubdivisionBase Object
nested_sub = city.subdivisions[0]
nested_sub.id
nested_sub.name
nested_sub.geonames_code
nested_sub.iso_code
nested_sub.type
nested_sub.admin_level
MacroregionBase Object
nested_macroregion = country.macroregions[0]
nested_macroregion.id
nested_macroregion.name
nested_macroregion.code
nested_macroregion.type
Performance
Caching
All registries and their indexes are lazy-loaded on first use, incurring a cold start cost on whichever call touches them first. Any registry's dataset and indexes can be pre-loaded with .force_cache() to avoid this during queries, or you can simply access the registry/method to trigger the lazy loading upfront.
A registry's dataset also loads the datasets it references, if they aren't cached yet. Countries load macroregions, subdivisions load countries, and cities load subdivisions and countries. Only those datasets load, not their indexes. The subdivisions and cities tables below exclude them, so a cold first call on cities also pays for the subdivisions and countries datasets.
Countries (281)
| Component | Load Time | Memory |
|---|---|---|
| Dataset | ~1ms | 215KB |
| Lookup index | < 1ms | 41KB |
| Filter index | < 1ms | 119KB |
| Search index | ~4ms | 381KB |
| Combined | ~6ms | 756KB |
Subdivisions (51,711)
| Component | Load Time | Memory |
|---|---|---|
| Dataset | ~85ms | 16.4MB |
| Lookup index | ~16ms | 4.3MB |
| Filter index | ~120ms | 13.8MB |
| Search index | ~62ms | 12.8MB |
| Combined | ~284ms | 47.4MB |
Cities (235,915)
⚠️ Memory-intensive. Fully caching cities and its indexes adds 114.4MB of memory. Calling
localis.cities.force_cache()loads all of it upfront. You can callcities.set_population_threshold(n)before first access as a lever to control the memory footprint.
| Component | Load Time | Memory |
|---|---|---|
| Dataset | ~341ms | 28.1MB |
| Lookup index | ~74ms | 1.8MB |
| Filter index | ~737ms | 48.5MB |
| Search index | ~150ms | 36.0MB |
| Combined | ~1.30s | 114.4MB |
At a threshold of 15,000, cities drops from 235,915 to 34,171 and memory drops from 114.4MB to 26.5MB.
Full Cache: ~1.58s load time, 162.5MB memory for all datasets and indexes
Benchmarks
Per-call query latency on warm caches, median (95th percentile), and fuzzy search accuracy on mangled/misspelled queries: how often the right entry appears in the top 10 results, and how often it's the top result.
| Registry | Get | Lookup | Filter | Search (p50/p95) | Search Accuracy (top 10) | Top Result |
|---|---|---|---|---|---|---|
| Countries | 0.0091ms (0.0125ms) | 0.0126ms (0.0166ms) | 0.018ms (0.0237ms) | 2.12ms (4.92ms) | 100.0% | 99.3% |
| Subdivisions | 0.0085ms (0.0105ms) | 0.0136ms (0.0168ms) | 0.0185ms (0.0363ms) | 3.33ms (5.77ms) | 97.0% | 85.2% |
| Cities | 0.0142ms (0.0181ms) | 0.012ms (0.0146ms) | 0.0238ms (0.0852ms) | 6.82ms (11.8ms) | 98.7% | 91.6% |
Accuracy tested on 5,000 mangled-query samples per registry; cities' search additionally includes city + admin1 context. Load times, memory and latency are generated by tests/analysis/footprint.py and tests/analysis/benchmarks.py, last measured on 11th Gen Intel(R) Core(TM) i7-1165G7 @ 2.80GHz with Python 3.14.7.
Data Sources
Data in this project is kept current monthly from the following sources:
- Countries
- Canonical: ISO 3166-1 data via Debian's iso-codes project
- Merged: ISO 3166-3 withdrawn/historic country codes, also via Debian's iso-codes project
- Merged: GeoNames
countryInfo.txt - Merged: Additional country aliases queried from Wikidata: English labels, alternative labels and short names
- Subdivisions
- Canonical: ISO 3166-2 data via Debian's iso-codes project
- Merged: GeoNames
admin1CodesASCII.txtandadmin2Codes.txt - Merged: Wikidata (ISO 3166-2 code ↔ GeoNames id)
- Merged: Additional subdivision aliases from GeoNames'
alternateNamesV2dump (filtered by Unicode CLDR's official-language data per country)
- Cities
- GeoNames
cities500.txtdataset
- GeoNames
- Macroregions
- Unicode CLDR territory containment and English territory names
docs/methodology.md is a complete, falsifiable account of how each dataset is built: the rules that combine these sources, how the results were validated, and where they are known to be wrong. unmerged_subdivisions.md lists every ISO subdivision currently without a GeoNames counterpart, regenerated on every ingest run.
Concurrency
Registries are safe to share across threads. get(), lookup(), filter(), search() and iteration only read shared data, and the first access that loads a dataset or index does so under the registry's lock, so threads reaching a cold registry together load it once.
Two settings change shared state for every thread: cities.set_population_threshold() and countries.set_include_historic(). Configure them before the registry is shared between threads, never while other threads are querying it.
Batch searching
localis doesn't parallelize batches for you, since the right approach depends on your Python build, memory budget and surrounding executor. On free-threaded Python (3.14t), a thread pool searches in parallel:
from concurrent.futures import ThreadPoolExecutor
import localis
localis.cities.force_cache() # load once, before the threads start
with ThreadPoolExecutor() as pool:
results = list(pool.map(localis.cities.search, queries))
On a standard Python build the same code is correct but runs one search at a time, because the GIL lets only one thread run Python code at once. To search in parallel there, use processes. Each worker loads its own copy of the data (up to 114.4MB for cities), so apply any settings in the worker's initializer, and search through a module-level function, since a registry itself can't be sent to a process:
from concurrent.futures import ProcessPoolExecutor
import localis
def init_worker():
localis.cities.set_population_threshold(15000) # repeat any setting the parent uses
def search_city(query):
return localis.cities.search(query)
if __name__ == "__main__": # workers import this module, so the pool only starts in the parent
with ProcessPoolExecutor(initializer=init_worker) as pool:
results = list(pool.map(search_city, queries, chunksize=500))
Data licensing
The shipped data is derived from these sources and remains subject to their licenses: ISO 3166 data via iso-codes (LGPL-2.1-or-later), GeoNames (CC BY 4.0), Wikidata (CC0), and Unicode CLDR (Unicode License v3).
Requirements
- Python 3.11+
rapidfuzz- Fast fuzzy string matchingunidecode- Unicode text normalization
License
Code: MIT. Data: subject to its sources' licenses, listed under Data licensing.
Why localis
localis began with some database cleanup. I found myself writing mountains of bespoke code to parse inconsistent, dirty data with pycountry, GeoNames and Wikidata to name a few (Google Places was not in the budget). When the pipeline was complete and the data cleaned, I realized this mountain of code could be useful for others who might need a reliable offline geo-data solution, so here we are! I hope you find it useful and please don't hesitate to contribute or report any issues.
Contributing
Pull requests welcome Report issues
Support this project:
Metadata
Release files for localis 2.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| localis-2.1.0.tar.gz | 21.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| localis-2.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 42.5 MB
Release files / localis-2.1.0.tar.gz
| Download URL | localis-2.1.0.tar.gz |
|---|---|
| Size | 21.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7f9f2080cc6820399a3c79a742e0ed37ae873cdc79a36a9972afc381f65b2f05
|
|
BLAKE2b-256 checksum How to use checksums |
61f8709f2523572f88bd040bcc732f8cdc92ed5dc88c8161d397f4418cac3491
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / localis-2.1.0-py3-none-any.whl
| Download URL | localis-2.1.0-py3-none-any.whl |
|---|---|
| Size | 21.4 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bcb629be25d5fa8f86382bef720258eb2c1a535ea03cdc787d7bd2827e45de79
|
|
BLAKE2b-256 checksum How to use checksums |
3088262a49c04f95f755f7f1a3c7eaf69687e3d929e6da69cff858a4190f59b2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log