localis
Fast, offline access to comprehensive data for countries, subdivisions, and cities. Built on ISO 3166 and GeoNames datasets (updated monthly) with support for exact lookups, filtering, and fuzzy search.
Features
- 🌍 281 countries (31 historic) sourced and merged from ISO 3166-1, ISO 3166-3, and GeoNames
- 🗺️ 51,711 subdivisions sourced and merged from ISO 3166-2 and GeoNames
- 🏙️ 235,914 cities sourced from GeoNames cities500.txt
- 🔍 Search Engine for typo-tolerant lookups with up to 89%+ accuracy
- 📌 Aliases - support for colloquial, historic and alternate names
Installation
pip install localis
Quick Start
import localis
# Countries
country = localis.countries.lookup("US")
print(country.name) # "United States"
# Subdivisions
state = localis.subdivisions.lookup("US-CA")
print(state.name) # "California"
# Fuzzy search
results = localis.countries.search("Austrlia") # Typo-tolerant
print(results[0][0].name) # "Australia"
Countries API
Get
import localis
# By localis ID
country = localis.countries.get(1)
Returns: Country object or None
Lookup
# By alpha-2 code
country = localis.countries.lookup("GB")
# By alpha-3 code
country = localis.countries.lookup("GBR")
# By numeric code
country = localis.countries.lookup(826)
Returns: Country object or None
Filter
# Exact name match (searches name, official_name, and aliases)
results = localis.countries.filter(name="Canada")
# General query across all fields
results = localis.countries.filter(name="United", limit=5)
Returns: list[Country]
Fuzzy Search
# Typo-tolerant search
results = localis.countries.search("Germny", limit=5)
for country, score in results:
print(f"{country.name}: {score}")
# Output:
# Germany: 0.951
# Guernsey: 0.714
# ...
Returns: list[tuple[Country, float]] - sorted by similarity score
Iteration
# Iterate over all countries
for country in localis.countries:
print(country.name)
# Get count
total = len(localis.countries)
Historic Countries
ISO 3166-3 withdrawn countries (Czechoslovakia, Serbia and Montenegro, Netherlands Antilles, and others) are included in the dataset but excluded from filter(), search(), and iteration by default.
localis.countries.include_historic # False
len(list(localis.countries)) # 250
# Include historic entries
localis.countries.set_include_historic(True)
len(list(localis.countries)) # 281
localis.countries.set_include_historic(False)
get() and lookup() always resolve historic entries regardless of the toggle. ISO reused alpha2/alpha3/numeric codes across different withdrawn countries over time (e.g. CS was both Czechoslovakia and, decades later, Serbia and Montenegro), so lookup() only resolves a historic entry by its unique alpha_4 withdrawal code, never by bare alpha2/alpha3/numeric:
Country Object
country = localis.countries.lookup("US")
country.id # Database ID
country.name # "United States"
country.official_name # "United States of America"
country.alpha2 # "US"
country.alpha3 # "USA"
country.geonames_id # 6252001
country.numeric # 840
country.aliases # tuple[str, ...] - Alternate names
country.flag # "🇺🇸" - Unicode flag emoji
country.historic # HistoricInfo | None - set only for withdrawn ISO 3166-3 countries
country = localis.countries.lookup("CSHH") # Czechoslovakia
country.historic.alpha_4 # "CSHH"
country.historic.withdrawal_date # "1993-01-01"
country.historic.comment # str | None
# Utility methods
country.to_dict() # Convert to dictionary
country.json() # Convert to JSON string
Subdivisions API
Get by ID
import localis
# By localis ID
subdivision = localis.subdivisions.get(1)
Returns: Subdivision object or None
Lookup by identifier
# By ISO code (country-subdivision)
subdivision = localis.subdivisions.lookup("US-CA")
# By GeoNames code
subdivision = localis.subdivisions.lookup("US.CA")
Returns: Subdivision object or None
Filter
# Exact name match
results = localis.subdivisions.filter(name="California")
# By subdivision type
results = localis.subdivisions.filter(type="state")
# By country
results = localis.subdivisions.filter(country="United States")
# By admin level (1 = states/provinces, 2 = counties/districts)
results = localis.subdivisions.filter(admin_level=1)
# Combine multiple filters (AND logic)
results = localis.subdivisions.filter(
country="US",
type="state",
limit=10
)
Returns: list[Subdivision]
Fuzzy Search
results = localis.subdivisions.search("Californa", limit=3)
for subdivision, score in results:
print(f"{subdivision.name}: {score}")
# California: 0.94
# Baja California: 0.8
# ...
Returns: list[tuple[Subdivision, float]]
Subdivision Object
subdivision = localis.subdivisions.lookup("US-CA")
subdivision.id # Database ID
subdivision.name # "California"
subdivision.geonames_code # "US.CA"
subdivision.iso_code # "US-CA"
subdivision.type # "State"
subdivision.admin_level # 1
subdivision.parent # SubdivisionBase | None - Parent subdivision
subdivision.country # CountryBase object
subdivision.aliases # tuple[str, ...] - Alternate names
# Utility methods
subdivision.to_dict() # Convert to dictionary
subdivision.json() # Convert to JSON string
Cities API
Get by ID
import localis
# By localis ID
city = localis.cities.get(1)
Returns: City object or None
Lookup by GeoNames id
# By GeoNames ID
city = localis.cities.lookup(5128581)
Returns: City object or None
Filter
# Exact name match
results = localis.cities.filter(name="Los Angeles")
# By country name or alpha2/alpha 3 code
results = localis.cities.filter(country="United States", limit=10)
# By subdivision name or ISO/GeoNames code
results = localis.cities.filter(subdivision="California", limit=10)
# Combine filters (AND logic)
results = localis.cities.filter(
country="US",
subdivision="California",
limit=20
)
Returns: list[City]
Fuzzy Search
results = localis.cities.search("Los Angelos", limit=5)
for city, score in results:
print(f"{city.name}, {city.country.name}: {score}")
Returns: list[tuple[City, float]] - sorted by similarity score
Population Threshold
# Narrow the cache and all indexes to cities with population >= 15000
localis.cities.set_population_threshold(15000)
# Check the current threshold
localis.cities.population_threshold # 15000
# Reset back to the full dataset
localis.cities.set_population_threshold(None)
cities is fully lazy-loaded, nothing is read from disk until first access. Call set_population_threshold() before that first access (before any .get(), .lookup(), .filter(), .search(), or .force_cache() call) so the registry only ever loads the narrowed dataset. Calling it after the cache or indexes are already built still works, but it invalidates them, so the next access rebuilds the caches from scratch at the new threshold.
Returns: None
City Object
city = localis.cities.lookup(5128581) # GeoNames ID
city.id # Database ID
city.geonames_id # 5128581
city.name # "New York"
city.subdivisions # list[SubdivisionBase] - ordered by admin_level ascending
city.country # CountryBase object
city.population # 8804190
city.lat # 40.71427
city.lng # -74.00597
# Utility methods
city.to_dict() # Convert to dictionary
city.json() # Convert to JSON string
Base Objects
Basic versions of country and subdivision when nested.
CountryBase Object
nested_country = subdivision.country
nested_country.id
nested_country.name
nested_country.alpha2
nested_country.alpha3
nested_country.geonames_id
SubdivisionBase Object
nested_sub = city.subdivisions[0]
nested_sub.id
nested_sub.name
nested_sub.geonames_code
nested_sub.iso_code
nested_sub.type
nested_sub.admin_level
Performance
Caching
All registries and their indexes are lazy-loaded on first use, incurring a cold start cost on whichever call touches them first. Any registry's dataset and indexes can be pre-loaded with .force_cache() to avoid this during queries, or you can simply access the registry/method to trigger the lazy loading upfront.
Countries (281)
| Component | Load Time | Memory |
|---|---|---|
| Dataset | < 1ms | 308KB |
| Lookup index | < 1ms | 28KB |
| Filter index | < 1ms | 164KB |
| Search index | ~2ms | 700KB |
| Combined | ~4ms | 1.2MB |
Subdivisions (51,711)
| Component | Load Time | Memory |
|---|---|---|
| Dataset | ~88ms | 27.2MB |
| Lookup index | ~13ms | 3.8MB |
| Filter index | ~125ms | 18.9MB |
| Search index | ~60ms | 11.2MB |
| Combined | ~287ms | 61.0MB |
Cities (235,914)
⚠️ Memory-intensive. Fully caching cities and its indexes adds 151.1MB of resident memory. Calling
localis.cities.force_cache()loads all of it upfront. You can callcities.set_population_threshold(n)before first access as a lever to control the memory footprint.
| Component | Load Time | Memory |
|---|---|---|
| Dataset | ~450ms | 60.4MB |
| Lookup index | ~63ms | 4KB |
| Filter index | ~580ms | 57.6MB |
| Search index | ~178ms | 33.1MB |
| Combined | ~1.27s | 151.1MB |
At a threshold of 15,000, cities drops from 235,914 to 34,171 and memory drops from 151.1MB to 31.1MB.
Full Cache: ~1.50s load time, 213.2MB memory for all datasets and indexes
Concurrency: localis is not yet thread-safe. Lazy loading can race on first access, and search() keeps per-query state on the shared index, so concurrent searches on the same registry can interfere with each other even after .force_cache(). Until thread safety lands, call .force_cache() up front and serialize searches on a shared registry (or give each thread its own process).
Benchmarks
Per-call query latency on warm caches, median (95th percentile), and fuzzy search accuracy on mangled/misspelled queries: how often the right entry appears in the top 10 results, and how often it's the top result.
| Registry | Get | Lookup | Filter | Search | Search Accuracy (top 10) | Top Result |
|---|---|---|---|---|---|---|
| Countries | 0.0021ms (0.0028ms) | 0.0058ms (0.0073ms) | 0.0093ms (0.0131ms) | 1.35ms (2.24ms) | 89.8% | 83.8% |
| Subdivisions | 0.0075ms (0.0101ms) | 0.012ms (0.0169ms) | 0.0153ms (0.0297ms) | 2.83ms (10.9ms) | 87.5% | 70.1% |
| Cities | 0.0142ms (0.0199ms) | 0.0118ms (0.0163ms) | 0.0216ms (0.0724ms) | 15.9ms (63.6ms) | 97.4% | 68.1% |
Accuracy tested on 5,000 mangled-query samples per registry; cities' search additionally includes city + admin1 context. Load times, memory and latency are generated by tests/analysis/footprint.py and tests/analysis/benchmarks.py, last measured on 11th Gen Intel(R) Core(TM) i7-1165G7 @ 2.80GHz with Python 3.14.4.
Data Sources
Data in this project is kept current monthly from the following sources:
- Countries
- Canonical: ISO 3166-1 data via Debian's iso-codes project
- Merged: ISO 3166-3 withdrawn/historic country codes, also via Debian's iso-codes project
- Merged: GeoNames
countryInfo.txt - Merged: Additional country aliases from a static Wikidata snapshot (not refreshed monthly)
- Subdivisions
- Canonical: ISO 3166-2 data via Debian's iso-codes project
- Merged: GeoNames
admin1CodesASCII.txtandadmin2Codes.txt - Merged: Wikidata crosswalk (ISO 3166-2 code ↔ GeoNames id) for unambiguous resolution ahead of fuzzy matching
- Merged: Additional subdivision aliases from GeoNames'
alternateNamesV2dump (filtered by Unicode CLDR's official-language data per country)
- Cities
- GeoNames
cities500.txtdataset
- GeoNames
docs/methodology.md is a complete, falsifiable account of how each dataset is built: the rules that combine these sources, how the results were validated, and where they are known to be wrong. unmerged_subdivisions.md lists every ISO subdivision currently without a GeoNames counterpart, regenerated on every ingest run.
Data licensing
The shipped data is derived from these sources and remains subject to their licenses: ISO 3166 data via iso-codes (LGPL-2.1-or-later), GeoNames (CC BY 4.0), Wikidata (CC0), and Unicode CLDR (Unicode License v3).
Requirements
- Python 3.11+
rapidfuzz- Fast fuzzy string matchingunidecode- Unicode text normalization
License
Code: MIT. Data: subject to its sources' licenses, listed under Data licensing.
Why localis
localis began with some database cleanup. I found myself writing mountains of bespoke code to parse inconsistent, dirty data with pycountry, GeoNames and Wikidata to name a few (Google Places was not in the budget). When the pipeline was complete and the data cleaned, I realized this mountain of code could be useful for others who might need a reliable offline geo-data solution, so here we are! I hope you find it useful and please don't hesitate to contribute or report any issues.
Contributing
Pull requests welcome Report issues
Support this project:
Metadata
Release files for localis 2.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| localis-2.0.0.tar.gz | 22.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| localis-2.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 44.6 MB
Release files / localis-2.0.0.tar.gz
| Download URL | localis-2.0.0.tar.gz |
|---|---|
| Size | 22.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3ad2e135abc735e3e626cf75582742613323d6652d3f5c62454158996a572e26
|
|
BLAKE2b-256 checksum How to use checksums |
bd04b4ebd91feb996c23cab7d3ac5a4d5eeb85fcf0ad4fd10d6bd03d515a3cfe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / localis-2.0.0-py3-none-any.whl
| Download URL | localis-2.0.0-py3-none-any.whl |
|---|---|
| Size | 22.4 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4c6293e9cafe472b32a73b5c8cb6b4f0aa002fa2782958d00eb498aa97b91ff0
|
|
BLAKE2b-256 checksum How to use checksums |
c004b861dfd091b564f50e22b5f3f24f812ec1c24a243c5a611330062f10c9a8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log