Parsernaam
Parsernaam uses two character-level LSTM classifiers to label a single token as
first or last, or a multi-token string as first_last or last_first. It
is useful when name fields were not collected separately and simple word-order
rules are inadequate.
These labels cannot represent every naming convention. Model scores are not calibrated guarantees, and errors and population imbalance in the training records can affect predictions. Do not use the output to infer ethnicity, citizenship, religion, gender, eligibility, or identity, or as the sole input to a consequential decision.
Installation
pip install parsernaam
Install the optional Gradio interface with:
pip install "parsernaam[web]"
Python API
import pandas as pd
from parsernaam import parse_names
names = pd.DataFrame(
{
"full_name": [
"Jan",
"Nicholas Turner",
"Nichols Richard",
"Kim Yeon",
]
},
index=pd.Index([10, 20, 30, 40], name="row_id"),
)
result = parse_names(names, names_col="full_name")
print(result[["full_name", "parsed_name"]])
parse_names returns a copy, preserves the input index and other columns, and
adds parsed_name. Each value contains the original string, one of the four
model labels, and its model score. Existing parsed_name values are replaced
without merge suffixes.
Invalid, blank, or non-Latin-only values receive the unknown label and a
score of 0.0. Input is normalized with Unicode accent removal, whitespace
collapse, and title-casing before batched inference. Case variants share the
same prediction and score. Apostrophes and hyphens are retained; unsupported
characters are removed. Long inputs are truncated at the model's manifest
sequence length. The result's name always preserves the original input.
Command line
The command-line interface uses Parquet for typed input and output:
parse_names input.parquet --output output.parquet --names-col full_name
The name column defaults to name, and the output path defaults to
output.parquet.
Model artifacts
The two SafeTensors weight files and string vocabulary are published
at gojiberries/parsernaam.
Parsernaam downloads them from an immutable Hugging Face commit and verifies
their SHA-256 hashes against the packaged model_manifest.json. Set
PARSERNAAM_MODEL_DIR to use an explicitly managed local copy. The Hugging
Face client honors its standard authentication configuration, including
HF_TOKEN.
Hub inference settings come from the packaged manifest. A manifest in an explicit local override supplies that local model's settings.
The published notebooks use early 2022 Florida voter registrations at Harvard Dataverse and a US Census surname list. They do not establish Indian training provenance. Their validation and test splits overlap training, so their reported quality is not a held-out estimate. Version 0.4 trains on FL registration fields only and evaluates geographic transfer on NC registration fields. The weights are hosted on Hugging Face; the package pins their exact revision.
Use local retrained artifacts with:
PARSERNAAM_MODEL_DIR=training/artifacts uv run parse_names input.parquet
A local model_manifest.json specifies architecture, encoding, sequence lengths,
labels, hashes, and evaluation provenance. Complete local manifests are checked
for missing or corrupted files before loading. Runtime inference uses the retrained packed models with padding 0 and unknown 1.
The original encoding is retained only in the historical evaluation harness.
Evaluation
The training pipeline documents the source fields, surname-disjoint FL splits, seeded record samples, and independent baselines. Published and fixed-weight FL comparisons remain contaminated; NC comparisons measure transfer between states and can contain shared name strings. Raw scores are evaluated for calibration, rather than treated as calibrated probabilities.
Accuracy on as-is input (NC is all caps):
| source | task | majority | frequency | published | fixed_inference | retrained |
|---|---|---|---|---|---|---|
| FL | ordering | 50.0% | 46.2% | 85.1% | 98.5% | 92.7% |
| FL | single_distinct_name | 49.8% | 45.2% | 79.4% | 86.5% | 81.6% |
| FL | single_record_weighted | 51.0% | 49.2% | 73.6% | 81.8% | 84.4% |
| NC | ordering | 50.0% | 87.1% | 53.2% | 98.4% | 96.2% |
| NC | single_distinct_name | 41.4% | 66.9% | 51.2% | 87.3% | 79.6% |
| NC | single_record_weighted | 50.0% | 84.9% | 54.2% | 80.9% | 90.9% |
Expected calibration error of raw scores (lower is better):
| source | task | majority | frequency | published | fixed_inference | retrained |
|---|---|---|---|---|---|---|
| FL | ordering | 0.000 | 0.433 | 0.130 | 0.008 | 0.020 |
| FL | single_distinct_name | 0.013 | 0.303 | 0.069 | 0.056 | 0.063 |
| FL | single_record_weighted | 0.001 | 0.399 | 0.096 | 0.129 | 0.060 |
| NC | ordering | 0.000 | 0.100 | 0.430 | 0.007 | 0.003 |
| NC | single_distinct_name | 0.097 | 0.066 | 0.224 | 0.087 | 0.084 |
| NC | single_record_weighted | 0.010 | 0.087 | 0.151 | 0.142 | 0.008 |
Release gates: case_invariance=True, case_violations=0, retrained_beats_both_baselines=True, fixed_beats_published=True.
The generated aggregate table is saved in training/reports/summary.md. The local HTML report includes score-band accuracy, ECE, support breakdowns, and 20 source-linked sample rows per task and state. The release uses the retrained models shown in the table. Record-weighted single-name accuracy improves over fixed original weights, but distinct-name and ordering accuracy are lower. On lookup-supported distinct NC names, frequency lookup scores 92.0% versus the retrained model's 81.3%. These tradeoffs matter when choosing whether the package suits a dataset. FL test surnames have no training surname counts by design, which explains the lookup's poor FL scores; NC measures transfer to a different state's records.
Development
uv sync --all-groups --all-extras
make ci
make docs
Authors
Rajashekar Chintalapati and Gaurav Sood
Related projects
- naamkaran generates synthetic name-like strings.
- ethnicolr is the canonical ethnicity-from-name package.
- pranaam estimates aggregate religion patterns from names.
License
Parsernaam is released under the MIT License.
Metadata
Release files for parsernaam 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| parsernaam-0.4.0.tar.gz | 14.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| parsernaam-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 31.2 kB
Release files / parsernaam-0.4.0.tar.gz
| Download URL | parsernaam-0.4.0.tar.gz |
|---|---|
| Size | 14.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
95fbcd2061a26a35f0c6f63adc786d1e31084c689dec98cb35df5639ff26d1f7
|
|
BLAKE2b-256 checksum How to use checksums |
0d350e9cda35ec1011a3a84ccd2b159286c286e98c889bce1e2712151d38e67c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / parsernaam-0.4.0-py3-none-any.whl
| Download URL | parsernaam-0.4.0-py3-none-any.whl |
|---|---|
| Size | 16.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5dc4531291833b33000b63ca67e247f120893682caaf2ce912b973543ef4b563
|
|
BLAKE2b-256 checksum How to use checksums |
9a9d636ac27fd98bb43831faf6916bcc71085f9a8e574b425e9df0b47badc103
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log