instate: state and language composition estimates for Indian surnames
Instate looks up electoral-roll surname shares across 35 states and union territories and estimates shares for unseen surnames with a calibrated character model. It derives language compositions by mixing state shares with Census 2011 mother-tongue shares.
The lookup contains 1,848,011 surname strings. Lookup and training retain only surname-state cells with at least three source occurrences; totals sum those retained cells. Most source rolls are from 2017, with Assam and Lakshadweep from 2026. The outputs describe name patterns, not an individual's residence, origin, or language.
Results follow the appeler inference contract, composition form: every row carries proportions that sum to one, explicit abstention with a machine-readable reason instead of a default distribution, and provenance columns identifying the exact artifacts used.
Installation
pip install instate
Usage
lookup_state_composition reports the electoral-roll shares for surnames in
the table and abstains on the rest:
import instate
result = instate.lookup_state_composition(["dhingra", "sood", "qzxv"])
result[
[
"surname",
"scored",
"abstention_reason",
"state_share_delhi",
"state_share_punjab",
"surname_record_count",
]
]
# surname scored abstention_reason state_share_delhi state_share_punjab surname_record_count
# dhingra True <NA> 0.530 0.231 7583
# sood True <NA> 0.194 0.364 29451
# qzxv False out-of-dictionary <NA> <NA> <NA>
estimate_state_composition runs the temperature-scaled BiLSTM for the same
quantity, including surnames the table has never seen:
result = instate.estimate_state_composition(["chintalapati"])
estimate_language_composition mixes state evidence with each state's
Census 2011 mother-tongue shares. By default it uses the lookup where the
surname is known and falls back to the model, recording which in a
language_basis column:
result = instate.estimate_language_composition(["sood", "chintalapati"])
result[["surname", "language_basis", "language_share_punjabi", "language_share_telugu"]]
DataFrame input uses the fleet signature: data first, then the column
name, with every option keyword-only.
import pandas as pd
frame = pd.DataFrame({"lastname": ["sharma", "patel"], "person_id": [1, 2]})
result = instate.lookup_state_composition(frame, "lastname")
Two reference lookups round out the API: lookup_state_official_languages
maps states to their official languages, and list_supported_states returns
the 35-state vocabulary.
What the outputs mean
The state shares' denominator is the surname's retained occurrences across included rolls. Cells with fewer than three occurrences are excluded before normalization from both lookup and training; their counts do not contribute to the published total. This is not a count of people in the current population. The model is trained so its softmax targets exactly that distribution, and its probabilities are temperature-scaled against held-out surnames, so the lookup and the estimate are two routes to one quantity.
The language composition is defined, not observed:
p(language | surname) = sum over states of
p(state | surname) x census mother-tongue share of the language in the state
The mother-tongue shares come from Census of India 2011 table C-16, with
Telangana aggregated from its ten 2011 districts and languages below a 1%
share in every state pooled into other
(builder, provenance and
hashes in the shipped manifest). Two caveats are part of the definition:
C-16 records mother tongue, not languages spoken, and the mixing assumes
language and surname are independent within a state, which understates
community-specific associations.
Known data weaknesses: Gujarat surnames remain noisy from OCR;
trailing-vowel spelling variants (Kannada patila,
Odia dasa) are merged into their canonical forms (patil, das).
Abstention
A surname the package cannot support gets abstained = True and a reason
from the contract's shared vocabulary (missing-name, no-letters,
unsupported-script, out-of-dictionary, insufficient-evidence), never a
default distribution. Supported input is romanized ASCII a to z; the
model additionally requires three supported characters.
Model and evaluation
The state model is a two-layer character-level bidirectional LSTM trained on the rebuilt 35-state data, with surnames assigned to deterministic disjoint train, validation, and test splits before training. Epoch selection uses the first 20,000 sorted names in the hash-assigned validation split; the best epoch is restored before saving. The other 165,007 validation names form a separate calibration set. Training and evaluation write manifests that bind the data bytes, checkpoint bytes, seed, and split membership; untouched-test evaluation refuses checkpoints without an eligible manifest (details).
35-state checkpoint metrics on the untouched test split, 185,232 surnames weighted by 61.4 million records:
| metric | value |
|---|---|
| modal state accuracy, top 1 / top 3 | 0.508 / 0.764 |
| record mass covered, top 1 / top 3 | 0.469 / 0.751 |
| record-weighted log loss, calibrated | 1.724 |
| top-1 confidence minus mass covered | -0.008 (0.074 before calibration) |
Calibration fits one temperature on those 165,007 reserved names against
each surname's retained empirical state distribution; the matching
instate_state_lstm_calibration.json records the temperature, objective,
and before/after metrics.
The matching model weights and lookup table download from a pinned Hugging
Face revision on first use. Downloads are cached and checked by SHA-256.
For offline use, set INSTATE_MODEL_DIR to a directory containing the
matching checkpoint, calibration JSON, and lookup Parquet.
Data
Sources: parsed electoral rolls and
source PDFs.
Census language shares rebuild from the pinned census downloads with
model_training/build_state_language_shares.py.
Coverage by state
The training rolls do not cover every state equally. Coverage below is the table's record weight divided by the state's electorate at the 2019 general election; a complete 2017 roll parse sits at 85 to 100 percent. Where a state is short, the model has fewer records to learn its surnames from, and a surname shared with a better-covered state is pulled toward that state.
| Coverage | States |
|---|---|
| 85 to 100 percent | Bihar, Odisha, Jharkhand, Goa, Tripura, Manipur, Maharashtra, Meghalaya, Haryana, Chandigarh, Puducherry, Punjab, Madhya Pradesh, Arunachal Pradesh, Mizoram, Uttarakhand, Sikkim, Tamil Nadu, Rajasthan, Uttar Pradesh, West Bengal, Himachal Pradesh |
| 80 to 85 percent | Nagaland, Daman and Diu, Telangana (English 2017 rolls, rebuilt in 3.1), Kerala |
| 55 to 70 percent | Dadra and Nagar Haveli, Andhra Pradesh, Delhi, Andaman and Nicobar Islands |
| under 55 percent | Gujarat (52 percent, OCR loss), Jammu and Kashmir and Ladakh (28 percent, Ladakh and the Jammu region only; the Urdu valley rolls are unparsed), Karnataka (15 percent, five northern districts only) |
| 2026 roll | Assam (the 2026 final roll, all 126 constituencies; 113 percent of the 2019 electorate) |
Lakshadweep uses 5,025 Latin surname selections from 57,618 active parsed
2026 entries; the shared lookup/training filters retain 3,312 occurrences across 381 strings.
This selective sample has much lower surname coverage than the complete
box parse. The model ranks Lakshadweep outside its top three for all 38
Lakshadweep-bearing test surnames (350 local record weight); the lookup
supplies direct evidence where a surname is present. The diagnostic is in
model_training/lakshadweep_2026_model_diagnostic.json. Chhattisgarh is not
in the vocabulary. Per-state sources,
build commands, and what each gap would take are in
model_training/prep_er_data/SOURCES.md.
Authors
Atul Dhingra, Gaurav Sood, and Rajashekar Chintalapati.
Contributor Code of Conduct
The project welcomes contributions from everyone! In fact, it depends on it. To maintain this welcoming atmosphere, and to collaborate in a fun and productive way, we expect contributors to the project to abide by the Contributor Code of Conduct.
License
The package is released under the MIT License.
Adjacent repositories
- appeler/naampy — Infer Sociodemographic Characteristics from Names Using Indian Electoral Rolls
- appeler/ethnicolr2 — Ethnicolr implementation with new models in pytorch
- appeler/parsernaam — AI name parsing. Predict first or last name using a DL model.
- appeler/ethnicolor — Race and Ethnicity based on name using data from census, voter reg. files, etc.
- appeler/ethnicolr — Predict Race and Ethnicity Based on the Sequence of Characters in a Name
Metadata
Release files for instate 3.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| instate-3.1.0.tar.gz | 22.5 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| instate-3.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 22.6 MB
Release files / instate-3.1.0.tar.gz
| Download URL | instate-3.1.0.tar.gz |
|---|---|
| Size | 22.5 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f11b45c4c0fd1dde785343fa5621bea918a9532376f219a587df56d9b0f443a8
|
|
BLAKE2b-256 checksum How to use checksums |
dcb8d0367cc5367176e51f89693904759f88f93575e1919fb8b9de979a718061
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.
Transparency logRelease files / instate-3.1.0-py3-none-any.whl
| Download URL | instate-3.1.0-py3-none-any.whl |
|---|---|
| Size | 52.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5acd3eb8a17604e56eb1b7b82813e77e08b0265602c10b37398f496ec9b13812
|
|
BLAKE2b-256 checksum How to use checksums |
d9ac772f73c4ac784a37bf12f4902ec84db150a0c430d59ca0d7a67b0cd7afb2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.
Transparency log