This release is a pre-release and may not be stable for production use.
localis
Fast, offline access to comprehensive data for countries, subdivisions, cities, the countries' macroregions, the currencies and languages they use, and the writing scripts languages are written in. Built on ISO 3166, ISO 4217, ISO 639-3, ISO 15924, GeoNames, Unicode CLDR and Wikidata datasets (updated monthly) with support for exact lookups, filtering, and fuzzy search.
⚠️ localis 3.0 is in beta.
pip install localisstill installs 2.1.x, which is no longer maintained; install the beta withpip install --pre localis. The CHANGELOG lists what changed from 2.1, and docs/versioning.md explains the beta.
Features
- 🌍 281 countries (31 historic) sourced and merged from ISO 3166-1, ISO 3166-3, GeoNames and Wikidata
- 🗺️ 51,711 subdivisions sourced and merged from ISO 3166-2 and GeoNames
- 🏙️ 235,970 cities sourced from GeoNames cities500.txt
- 🌐 34 macroregions (5 regions, 23 subregions, 6 groupings) sourced from Unicode CLDR, with every current country placed in them
- 💱 178 currencies and funds from ISO 4217, with each current country's legal tender from Unicode CLDR
- ✍️ 226 scripts from ISO 15924, with Unicode CLDR's English script names as aliases
- 🗣️ 7,923 languages from ISO 639-3, with Unicode CLDR's scripts for each and each current country's official languages
- 🔍 Typo-tolerant search: with a typo in the query, the intended record ranks first for 99.3% of countries, 93.9% of subdivisions and 93.6% of cities
- 📌 Aliases: support for colloquial, historic and alternate names
Installation
pip install --pre localis
localis follows Semantic Versioning; docs/versioning.md says what each release can change and which releases are supported.
Requirements
- Python 3.11+
rapidfuzz- Fast fuzzy string matching
Quick Start
import localis
# Exact lookups by code
country = localis.countries.lookup("US")
print(country.name) # "United States"
state = localis.subdivisions.lookup("US-CA")
print(state.name) # "California"
# Filters: exact matches, combined with AND
cities = localis.cities.filter(country="US", subdivision="California", limit=10)
# Typo-tolerant search: (result, score) pairs, best first
results = localis.countries.search("Austrlia")
print(results[0][0].name) # "Australia"
Querying
Each dataset is represented by its registry: countries, subdivisions, cities, macroregions, currencies, scripts and languages, sharing a common query API.
| Registry | get |
lookup |
filter |
search |
Iteration |
|---|---|---|---|---|---|
countries |
✓ | ✓ | ✓ | ✓ | ✓ |
subdivisions |
✓ | ✓ | ✓ | ✓ | ✓ |
cities |
✓ | ✓ | ✓ | ✓ | ✓ |
macroregions |
✓ | ✓ | ✓ | ||
currencies |
✓ | ✓ | ✓ | ✓ | ✓ |
scripts |
✓ | ✓ | ✓ | ✓ | ✓ |
languages |
✓ | ✓ | ✓ | ✓ | ✓ |
Lookups, filters and search ignore case and accents, so "sao paulo" finds São Paulo and "strasse" finds Straße.
get
country = localis.countries.get(1)
subdivision = localis.subdivisions.get(1)
city = localis.cities.get(1)
macroregion = localis.macroregions.get(1)
currency = localis.currencies.get(1)
script = localis.scripts.get(1)
language = localis.languages.get(1)
Returns: the entity with that localis ID, or None
ℹ️ IDs are only valid within the installed version; to store a reference, use key.
lookup
Resolves a single record by an identifier other than its localis ID.
| Registry | Identifiers |
|---|---|
countries |
alpha-2 ("GB"), alpha-3 ("GBR"), numeric (826), a historic entry only by its alpha_4 ("CSHH", see Historic Countries) |
subdivisions |
ISO 3166-2 code ("US-CA"), GeoNames code ("US.CA") |
cities |
GeoNames ID (5128581) |
macroregions |
code ("155", "EU"), name ("Western Europe") |
currencies |
alpha-3 ("EUR"), numeric (978) |
scripts |
alpha-4 ("Cyrl"), numeric (220) |
languages |
ISO 639-3 ("deu"), ISO 639-1 ("de"), ISO 639-2/B ("ger") |
country = localis.countries.lookup("GB")
subdivision = localis.subdivisions.lookup("US-CA")
city = localis.cities.lookup(5128581)
macroregion = localis.macroregions.lookup("EU")
currency = localis.currencies.lookup("EUR")
script = localis.scripts.lookup("Cyrl")
language = localis.languages.lookup("de")
Returns: the entity, or None. A key that isn't a string or an int raises TypeError.
ℹ️
lookup()matches identifiers only. Common abbreviations that aren't ISO codes, such as "UK" for the United Kingdom, are found byfilter(name=...)andsearch(). M49 codes are zero-padded strings, somacroregions.lookup("009")finds Oceania andlookup(9)finds nothing.
filter
Exact matches on any value a field indexes. Using multiple fields combines the conditions with a logical AND.
| Registry | Fields |
|---|---|
countries |
name (name, official name, common name or alias), macroregion (a region, subregion or grouping, by name or code), currency (by name or alpha-3), language (an official language, by name or ISO 639 code) |
subdivisions |
name (name or alias), type, country (name, common name, alpha-2, alpha-3 or numeric, as 76 or "076"), admin_level (0 = non-administrative groupings, 1 = states/provinces, 2 = counties/districts, 3 = divisions below those) |
cities |
name, country (name, common name, alpha-2 or alpha-3), subdivision (any subdivision in the city's chain, by name, ISO code or its suffix ("CA"), or GeoNames code) |
currencies |
name |
scripts |
name (name or alias) |
languages |
name (name, inverted name or alias), scope, type, script (by alpha-4, name or alias, primary or secondary) |
localis.countries.filter(name="UK") # aliases and abbreviations match too
localis.countries.filter(macroregion="Western Europe")
localis.countries.filter(currency="EUR") # countries whose legal tender includes the euro
localis.countries.filter(language="fr") # countries where French has an official status
localis.languages.filter(script="Cyrl", type="living")
localis.subdivisions.filter(country="US", type="state")
localis.subdivisions.filter(admin_level=1, limit=10)
localis.cities.filter(country="US", subdivision="California", limit=20)
# Records with no value in a field
localis.subdivisions.filter(type=localis.MISSING) # GeoNames-only subdivisions, which have no ISO type
localis.cities.filter(country="US", subdivision=localis.MISSING) # US cities linked to no subdivision
localis.countries.filter(macroregion=localis.MISSING) # historic countries placed in none (with include_historic set)
localis.countries.filter(currency=localis.MISSING) # countries with no legal tender, such as Antarctica
localis.languages.filter(script=localis.MISSING) # languages CLDR lists no script for
Pass localis.MISSING to match records with no value in a field (None ignores the field). A field the registry doesn't have raises TypeError, as does a call with no field.
Returns: a list of entities sorted by name. limit defaults to every match and must be at least 1.
search
for country, score in localis.countries.search("Germny", limit=5):
print(f"{country.name}: {score:.2f}")
# Germany first, then weaker matches such as Guernsey
localis.subdivisions.search("Californa")
localis.cities.search("Springfeld, Illinois") # context after the name narrows the match
localis.cities.search("Springfield", population_sort=True) # the best matches, largest city first
localis.currencies.search("Swiss Frank")
localis.scripts.search("Devanagri")
localis.languages.search("Portugese")
A subdivision or city query can add context after the name: a subdivision's parent or country, or a city's first-level subdivision or country.
Returns: a list of (entity, score) pairs, best match first. Each score runs from 0 to 1, higher is better. limit defaults to 10 and must be at least 1.
Iteration and len()
for country in localis.countries:
print(country.name)
total = len(localis.subdivisions)
ℹ️ Historic countries are left out of
filter(),search(), iteration andlen()unless included (see Historic Countries), and a population threshold narrows every cities query (see Population Threshold).
Entities
Results are typed dataclasses, listed field by field under each registry below. Each has to_dict() and json(), str() gives its JSON, and key is its stable reference (see key below). A record nested in another, such as subdivision.country, city.subdivisions, country.macroregions or country.currencies, is its base form (CountryBase, SubdivisionBase, MacroregionBase, CurrencyBase), which keeps the fields marked Base in those tables. A nested record that carries facts about the relationship, such as a country's languages or a language's scripts, is its base form plus those facts (CountryLanguage, LanguageScript).
Every returned entity is built fresh; yours to mutate freely. Entities are also hashable, so they can be used in sets and as dict keys.
Every entity type, their base Entity, Missing and the registry classes (CountryRegistry and the rest) import from localis for type annotations.
country = localis.countries.lookup("US")
country.to_dict() # dict of every field
country.json() # the same, as a JSON string
localis.subdivisions.lookup("US-CA").country.alpha3 # "USA", from the nested CountryBase
key
localis IDs are not stable across builds (localis.__version__ gives the installed one). Use the key attribute to reliably reference entities instead, which returns the entity's stable lookup identifier. A nested record also carries its key.
# 5128581, safe to persist
saved = city.key
# the same city, in this version or a later one
city = localis.cities.lookup(saved)
| Registry | key |
|---|---|
countries |
alpha2, or historic.alpha_4 for a historic entry |
subdivisions |
iso_code, or geonames_code for a subdivision ISO doesn't list |
cities |
geonames_id |
macroregions |
code |
currencies |
alpha3 |
scripts |
alpha4 |
languages |
alpha3 (ISO 639-3) |
Countries
Historic Countries
ISO 3166-3 withdrawn countries are included in the dataset but excluded from filter(), search(), iteration and len() by default. The countries registry exposes set_include_historic() to toggle them on or off.
localis.countries.include_historic # False
len(localis.countries) # current countries only
# Include historic entries
localis.countries.set_include_historic(True)
len(localis.countries) # current and historic
localis.countries.set_include_historic(False)
len(localis.countries) is 250 by default and 281 with historic entries included.
⚠️ The toggle applies to every thread using
localis.countries, so set it before sharing the registry between threads (see Concurrency).
get() and lookup() always find historic entries; lookup() finds them only by alpha_4.
localis.countries.lookup("CSHH") # Czechoslovakia
localis.countries.lookup("CS") # None: a reused code never resolves a historic entry
Country Object
A ✓ under Base marks a field the nested CountryBase also has.
| Field | Type | Example ("US") |
Notes | Base |
|---|---|---|---|---|
id |
int |
localis ID, valid within this version | ✓ | |
key |
str |
"US" |
stable reference to store; alpha_4 for a historic entry |
✓ |
name |
str |
"United States" |
ISO 3166-1 name, as published | ✓ |
official_name |
str | None |
"United States of America" |
ISO 3166-1 official name; None where ISO has none |
|
common_name |
str | None |
None |
Debian iso-codes' everyday name where it differs, such as "South Korea" for "Korea, Republic of" | |
alpha2 |
str |
"US" |
✓ | |
alpha3 |
str | None |
"USA" |
✓ | |
numeric |
int | None |
840 |
ISO 3166-1 numeric code; None for Kosovo, which has no ISO assignment |
|
geonames_id |
int | None |
6252001 |
✓ | |
aliases |
list[str] |
alternate names from GeoNames and Wikidata | ||
flag |
str | None |
"🇺🇸" |
Unicode flag emoji | |
historic |
HistoricInfo | None |
None |
set only for withdrawn ISO 3166-3 countries | |
macroregions |
list[MacroregionBase] |
[Americas, Northern America] | CLDR path, region then subregion; [] for most historic countries |
|
groupings |
list[MacroregionBase] |
[North America, United Nations] | CLDR groupings the country belongs to | |
currencies |
list[CurrencyBase] |
[US Dollar] | legal tender in use, per CLDR, in CLDR's order; [] for historic countries |
|
languages |
list[CountryLanguage] |
[English, Spanish, Hawaiian] | languages with an official status, per CLDR, by population share; [] for historic countries (see CountryLanguage) |
HistoricInfo
| Field | Type | Example ("CSHH") |
Notes |
|---|---|---|---|
alpha_4 |
str |
"CSHH" |
ISO 3166-3 withdrawal code, unique to the entry |
withdrawal_date |
str |
"1993-01-01" |
|
comment |
str | None |
ISO's comment |
Subdivisions
Subdivision Object
A ✓ under Base marks a field the nested SubdivisionBase also has.
| Field | Type | Example ("US-CA") |
Notes | Base |
|---|---|---|---|---|
id |
int |
localis ID, valid within this version | ✓ | |
key |
str |
"US-CA" |
stable reference to store; the GeoNames code where ISO doesn't list the subdivision | ✓ |
name |
str |
"California" |
ISO 3166-2 name where ISO lists the subdivision, otherwise GeoNames' | ✓ |
iso_code |
str | None |
"US-CA" |
None for a GeoNames-only subdivision |
✓ |
geonames_code |
str | None |
"US.CA" |
GeoNames admin code; None for an ISO-only subdivision |
✓ |
geonames_id |
int | None |
5332921 |
None for an ISO-only subdivision |
✓ |
type |
str | None |
"State" |
ISO type; None for a GeoNames-only subdivision |
✓ |
admin_level |
int |
1 |
0 = non-administrative grouping, 1 = top-level, 2 = second-level, 3 = below that | ✓ |
parent |
SubdivisionBase | None |
None |
the subdivision it sits in | |
country |
CountryBase |
United States | ||
aliases |
list[str] |
alternate names |
Cities
Population Threshold
# Narrow the cache and all indexes to cities with population >= 15000
localis.cities.set_population_threshold(15000)
# Check the current threshold
localis.cities.population_threshold # 15000
# Reset back to the full dataset
localis.cities.set_population_threshold(None)
Call cities.set_population_threshold() before first access so the registry only ever caches the narrowed dataset. Calling it after the cache or indexes are already built still works, but it invalidates them, so the next access rebuilds everything from scratch at the new threshold, paying the cache tax twice.
⚠️ The threshold applies to every thread using
localis.cities, so set it before sharing the registry between threads (see Concurrency).
City Object
| Field | Type | Example (5128581) |
Notes |
|---|---|---|---|
id |
int |
localis ID, valid within this version | |
key |
int |
5128581 |
stable reference to store, the GeoNames ID |
geonames_id |
int |
5128581 |
|
name |
str |
"New York" |
GeoNames' name, or its ASCII form where the name is in another script |
subdivisions |
list[SubdivisionBase] |
the city's full subdivision chain, ordered by admin_level ascending |
|
country |
CountryBase |
United States | |
population |
int |
GeoNames population; 0 where unknown |
|
lat |
float |
40.71427 |
|
lng |
float |
-74.00597 |
Macroregions
Macroregion Object
A ✓ under Base marks a field the nested MacroregionBase also has.
| Field | Type | Example ("155") |
Notes | Base |
|---|---|---|---|---|
id |
int |
localis ID, valid within this version | ✓ | |
key |
str |
"155" |
stable reference to store, the code | ✓ |
name |
str |
"Western Europe" |
CLDR English name | ✓ |
code |
str |
"155" |
M49 numeric code as a zero-padded string, or CLDR's letter code ("QO", "EU") |
✓ |
type |
MacroregionType |
"subregion" |
"region", "subregion" or "grouping" |
✓ |
parent |
MacroregionBase | None |
Europe | a subregion's region, or the region CLDR files a grouping under |
Currencies
Every ISO 4217 code ships as published, funds, precious metals and the testing codes included. Each current country's legal tender comes from Unicode CLDR (see Country Object).
Currency Object
A ✓ under Base marks a field the nested CurrencyBase also has.
| Field | Type | Example ("EUR") |
Notes | Base |
|---|---|---|---|---|
id |
int |
localis ID, valid within this version | ✓ | |
key |
str |
"EUR" |
stable reference to store, the alpha-3 | ✓ |
name |
str |
"Euro" |
ISO 4217 name, as published | ✓ |
alpha3 |
str |
"EUR" |
✓ | |
numeric |
int | None |
978 |
ISO 4217 numeric code |
Scripts
Every ISO 15924 code ships as ISO publishes it, including the special codes ("Zyyy" undetermined, "Zxxx" unwritten, "Zmth" mathematical notation) and the two entries marking the private-use range, "Qaaa" (start) and "Qabx" (end). ISO's names often carry other names in parentheses, such as "Han (Hanzi, Kanji, Hanja)"; Unicode CLDR's English names for the same code ("Han", "Simplified Han") are its aliases, so both are found by filter(name=...) and search().
Script Object
A ✓ under Base marks a field the nested ScriptBase also has.
| Field | Type | Example ("Deva") |
Notes | Base |
|---|---|---|---|---|
id |
int |
localis ID, valid within this version | ✓ | |
key |
str |
"Deva" |
stable reference to store, the alpha-4 | ✓ |
name |
str |
"Devanagari (Nagari)" |
ISO 15924 name, as published | ✓ |
alpha4 |
str |
"Deva" |
✓ | |
numeric |
int | None |
315 |
ISO 15924 numeric code | |
aliases |
list[str] |
["Devanagari"] | Unicode CLDR's English names for the code, where they differ from ISO's |
Languages
Every ISO 639-3 code ships as ISO publishes it: living, extinct, historical and constructed languages, macrolanguages such as Arabic and Chinese, and the special codes ("und" undetermined, "mul" multiple, "zxx" no linguistic content). A language's scripts come from Unicode CLDR, which covers 811 of them; the rest have none.
german = localis.languages.lookup("de")
german.scripts # [LanguageScript(alpha4="Latn", secondary=False, ...)]
for language in localis.countries.lookup("HK").languages:
print(language.name, language.status, language.population_percent, language.script)
Language Object
A ✓ under Base marks a field the nested LanguageBase also has.
| Field | Type | Example ("deu") |
Notes | Base |
|---|---|---|---|---|
id |
int |
localis ID, valid within this version | ✓ | |
key |
str |
"deu" |
stable reference to store, the ISO 639-3 code | ✓ |
name |
str |
"German" |
ISO 639-3 name, as published | ✓ |
alpha3 |
str |
"deu" |
ISO 639-3 code | ✓ |
alpha2 |
str | None |
"de" |
ISO 639-1 code, for the 184 languages that have one | ✓ |
bibliographic |
str | None |
"ger" |
ISO 639-2/B code, where it differs from the 639-3 one | |
scope |
LanguageScope |
"individual" |
"individual", "macrolanguage" or "special" |
|
type |
LanguageType |
"living" |
"living", "extinct", "historical", "constructed" or "special" |
|
inverted_name |
str | None |
None |
ISO's name with the qualifier moved last, such as "Arabic, Algerian Saharan" | |
aliases |
list[str] |
["Austrian German", ...] | Unicode CLDR's English names for the language and its regional and script forms, where they differ from ISO's | |
scripts |
list[LanguageScript] |
[Latin] | per CLDR, primary scripts first; see LanguageScript |
LanguageScript
A script as one language uses it: every ScriptBase field plus secondary.
| Field | Type | Notes |
|---|---|---|
secondary |
bool |
CLDR's rule: True when the language isn't a modern language or the script isn't a modern script, such as Arabic written in Syriac or anything in Sanskrit |
languages.filter(script=...) matches primary and secondary scripts alike.
CountryLanguage
A language as one country recognizes it: every LanguageBase field plus three from Unicode CLDR. A country lists one per language and script CLDR gives an official status, ordered by population_percent.
| Field | Type | Example (Hong Kong's Chinese) | Notes |
|---|---|---|---|
status |
LanguageStatus |
"official" |
"official", "regional" (official in part of the country) or "de_facto" (official in practice, such as English in the US) |
population_percent |
float | None |
95.0 |
CLDR's estimate of the population using it; shares overlap, since people use several languages |
script |
ScriptBase | None |
Han (Traditional variant) | the script CLDR gives the status for, so a language can appear once per script; None where CLDR names none |
Performance
Caching
All registries and their indexes are lazy-loaded on first use, incurring a cold start cost on whichever call touches them first. Any registry's dataset and indexes can be pre-loaded with .force_cache() to avoid this during queries, or you can simply access the registry/method to trigger the lazy loading upfront.
Each component is shown as load time / memory, measured for that registry alone.
| Registry | Records | Dataset | Lookup index | Filter index | Search index | Combined |
|---|---|---|---|---|---|---|
| Macroregions | 34 | < 1ms / 7KB | < 1ms / 6KB | n/a | n/a | < 1ms / 13KB |
| Currencies | 178 | < 1ms / 29KB | < 1ms / 15KB | < 1ms / 20KB | < 1ms / 181KB | ~1ms / 245KB |
| Scripts | 226 | < 1ms / 48KB | < 1ms / 19KB | < 1ms / 29KB | ~1ms / 231KB | ~2ms / 327KB |
| Languages | 7,923 | ~11ms / 1.9MB | ~2ms / 589KB | ~7ms / 1.4MB | ~10ms / 1.5MB | ~31ms / 5.4MB |
| Countries | 281 | ~1ms / 369KB | < 1ms / 41KB | ~2ms / 229KB | ~2ms / 474KB | ~5ms / 1.1MB |
| Subdivisions | 51,711 | ~78ms / 14.5MB | ~16ms / 4.5MB | ~62ms / 12.1MB | ~58ms / 12.4MB | ~215ms / 43.5MB |
| Cities | 235,970 | ~341ms / 29.5MB | ~75ms / 1.9MB | ~282ms / 51.5MB | ~195ms / 44.0MB | ~894ms / 126.9MB |
| Total | 296,323 | ~433ms / 46.3MB | ~94ms / 7.1MB | ~354ms / 65.3MB | ~268ms / 58.7MB | ~1.15s / 177.5MB |
ℹ️ A registry also loads the datasets it references (without their indexes) so a cold first call on cities also loads countries and subdivisions, for example.
⚠️ Fully caching cities and its indexes adds 126.9MB of memory. Calling
localis.cities.force_cache()loads all of it upfront. You can callcities.set_population_threshold(n)before first access as a lever to control the memory footprint. At a threshold of 15,000, cities drops from 235,970 to 34,172 and memory drops from 126.9MB to 29.8MB.
Search Benchmarks
get(), lookup() and filter() return in microseconds on warm caches.
| Registry | Latency (p50 / p95) | Queries | Accuracy (top 10) | Top Result |
|---|---|---|---|---|
| Countries | 1.39ms / 4.23ms | 430 | 100.0% | 99.3% |
| Subdivisions | 2.6ms / 4.72ms | 7,005 | 98.0% | 93.9% |
| Cities | 6.4ms / 11.4ms | 5,000 | 98.5% | 93.6% |
Accuracy is tested on mangled names and aliases of up to 5,000 sampled records per registry; cities' queries add the city's admin1. A record the query can't tell apart from the sampled one, sharing its name (and a city's admin1), counts as a hit. Measured on 11th Gen Intel(R) Core(TM) i7-1165G7 @ 2.80GHz with Python 3.14.7.
Concurrency
Registries are safe to share across threads, and a cold registry loads once even when several threads reach it together.
⚠️ Two settings change shared state for every thread:
cities.set_population_threshold()andcountries.set_include_historic(). Configure them before the registry is shared between threads, never while other threads are querying it.
Batch searching
On free-threaded Python (3.14t), a thread pool searches in parallel:
from concurrent.futures import ThreadPoolExecutor
import localis
localis.cities.force_cache() # load once, before the threads start
with ThreadPoolExecutor() as pool:
results = list(pool.map(localis.cities.search, queries))
On a standard build, use processes. Each worker loads its own copy of the data (up to 126.9MB for cities), so apply settings in the worker's initializer:
from concurrent.futures import ProcessPoolExecutor
import localis
def init_worker():
localis.cities.set_population_threshold(15000) # repeat any setting the parent uses
def search_city(query):
return localis.cities.search(query)
if __name__ == "__main__": # workers import this module, so the pool only starts in the parent
with ProcessPoolExecutor(initializer=init_worker) as pool:
results = list(pool.map(search_city, queries, chunksize=500))
Data Sources
Data in this project is kept current monthly from the following sources:
- Countries
- Canonical: ISO 3166-1 data via Debian's iso-codes project
- Merged: ISO 3166-3 withdrawn/historic country codes, also via Debian's iso-codes project
- Merged: GeoNames
countryInfo.txt - Merged: Additional country aliases queried from Wikidata: English labels, alternative labels and short names
- Subdivisions
- Canonical: ISO 3166-2 data via Debian's iso-codes project
- Merged: GeoNames
admin1CodesASCII.txtandadmin2Codes.txt - Merged: Wikidata (ISO 3166-2 code ↔ GeoNames id)
- Merged: Additional subdivision aliases from GeoNames'
alternateNamesV2dump (Latin-script names only, filtered by Unicode CLDR's official-language data per country)
- Cities
- GeoNames
cities500.txtdataset
- GeoNames
- Macroregions
- Unicode CLDR territory containment and English territory names
- Currencies
- Canonical: ISO 4217 data via Debian's iso-codes project
- Merged: each country's legal tender from Unicode CLDR's currency data
- Scripts
- Canonical: ISO 15924 data via Debian's iso-codes project
- Merged: Unicode CLDR's English script names, as aliases
- Languages
- Canonical: ISO 639-3 data via Debian's iso-codes project
- Merged: Unicode CLDR's English language names as aliases, each language's scripts, and each country's official languages
Names ship in Latin script.
docs/methodology.md is a complete, falsifiable account of how each dataset is built: the rules that combine these sources, how the results were validated, and where they are known to be wrong. unmerged_subdivisions.md lists every ISO subdivision currently without a GeoNames counterpart.
Data licensing
The shipped data is derived from these sources, modified by localis's ingest pipeline, and remains under their licenses: ISO 3166, ISO 4217, ISO 639-3 and ISO 15924 data via iso-codes (LGPL-2.1-or-later), GeoNames (CC BY 4.0), Unicode CLDR (Unicode License v3) and Wikidata (CC0). src/localis/data/NOTICE, which ships with the data, attributes each source, and the full license texts are in LICENSES/ and in the wheel's metadata. If you redistribute the data, keep that notice and those licenses with it.
License
Code: MIT (LICENSE).
Data: its sources' licenses, listed under Data licensing.
The package's license expression is MIT AND LGPL-2.1-or-later AND CC-BY-4.0 AND Unicode-3.0 AND CC0-1.0.
History
localis began with some database cleanup. I found myself writing mountains of bespoke code to parse inconsistent, dirty data while trying to reconcile pycountry, GeoNames and Wikidata to name a few (Google Places was not in the budget). When the pipeline was complete and the data finally cleaned, I realized this mountain of code could be useful for others who might need a reliable offline solution, so here we are! Over the past few years I've taken great care to build a robust, reliable dataset, wrapped in a simple, performant interface. I hope you find it useful and please don't hesitate to contribute or report any issues.
Contributing
Support this project:
Contributors
- @SchubmannM: performance work in #8 that inspired localis's
array('I')index storage and lazy-loaded registries
Metadata
Release files for localis 3.0.0b1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| localis-3.0.0b1.tar.gz | 24.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| localis-3.0.0b1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 49.3 MB
Release files / localis-3.0.0b1.tar.gz
| Download URL | localis-3.0.0b1.tar.gz |
|---|---|
| Size | 24.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0259dbee3d9ef9de6ae5d629e237d0af24d5e6ebe20203aa002b027645597e52
|
|
BLAKE2b-256 checksum How to use checksums |
522059d0fa8c35ee1b94521d27bcb63d846792c1164507d47daad6d2cbcadb3a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / localis-3.0.0b1-py3-none-any.whl
| Download URL | localis-3.0.0b1-py3-none-any.whl |
|---|---|
| Size | 24.7 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d3cc54968ded2ebb0f7d4e4c73f68f851140c5fe8f0fb5e32def9927fbd62a28
|
|
BLAKE2b-256 checksum How to use checksums |
f9406774693563b4b60d20e6879effa78f320e30e09e768dda65e0317699c3f1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log