Parsernaam
Parsernaam uses two character-level LSTM classifiers to label a single token as
first or last, or a multi-token string as first_last or last_first. It
is useful when name fields were not collected separately and simple word-order
rules are inadequate.
These labels cannot represent every naming convention. Model scores are not calibrated guarantees, and errors and population imbalance in the training records can affect predictions. Do not use the output to infer ethnicity, citizenship, religion, gender, eligibility, or identity, or as the sole input to a consequential decision.
Installation
pip install parsernaam
Install the optional Gradio interface with:
pip install "parsernaam[web]"
Python API
import pandas as pd
from parsernaam import parse_names
names = pd.DataFrame(
{
"full_name": [
"Jan",
"Nicholas Turner",
"Nichols Richard",
"Kim Yeon",
]
},
index=pd.Index([10, 20, 30, 40], name="row_id"),
)
result = parse_names(names, names_col="full_name")
print(result[["full_name", "parsed_name"]])
parse_names returns a copy, preserves the input index and other columns, and
adds parsed_name. Each value contains the original string, one of the four
model labels, and its model score. Existing parsed_name values are replaced
without merge suffixes.
Invalid or blank values receive the unknown label and a score of 0.0.
Command line
The command-line interface uses Parquet for typed input and output:
parse_names input.parquet --output output.parquet --names-col full_name
The name column defaults to name, and the output path defaults to
output.parquet.
Model artifacts
The two PyTorch state dictionaries and non-null string vocabulary are published
at gojiberries/parsernaam.
Parsernaam downloads them from an immutable Hugging Face commit and verifies
their SHA-256 hashes against the packaged model_manifest.json. Set
PARSERNAAM_MODEL_DIR to use an explicitly managed local copy. The Hugging
Face client honors its standard authentication configuration, including
HF_TOKEN.
The repository documentation describes training records derived from Indian and United States voter registrations and cites the early 2022 Florida voter registration data at Harvard Dataverse. A complete row-level training manifest is not available, so use the models for exploration rather than population claims.
Development
uv sync --all-groups --all-extras
make ci
make docs
Authors
Rajashekar Chintalapati and Gaurav Sood
Related projects
- naamkaran generates synthetic name-like strings.
- ethnicolr is the canonical ethnicity-from-name package.
- pranaam estimates aggregate religion patterns from names.
License
Parsernaam is released under the MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file parsernaam-0.3.0.tar.gz.
File metadata
- Download URL: parsernaam-0.3.0.tar.gz
- Upload date:
- Size: 10.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
23c10c6c8379a4e3c4fb3741774139b326216a814a9cd61dbb7079d6e1235267
|
|
| MD5 |
d5c64a7e60dc1fba0e30ec00c0ffa22d
|
|
| BLAKE2b-256 |
f45d4a013ac912562d06fa9a952a1080041f61f7226d6cff264e13b846b91dee
|
Provenance
The following attestation bundles were made for parsernaam-0.3.0.tar.gz:
Publisher:
release.yml on appeler/parsernaam
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
parsernaam-0.3.0.tar.gz -
Subject digest:
23c10c6c8379a4e3c4fb3741774139b326216a814a9cd61dbb7079d6e1235267 - Sigstore transparency entry: 2500454402
- Sigstore integration time:
-
Permalink:
appeler/parsernaam@e2c618e18a3ad5ec42d3051bcb209b13094b1a58 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/appeler
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@e2c618e18a3ad5ec42d3051bcb209b13094b1a58 -
Trigger Event:
push
-
Statement type:
File details
Details for the file parsernaam-0.3.0-py3-none-any.whl.
File metadata
- Download URL: parsernaam-0.3.0-py3-none-any.whl
- Upload date:
- Size: 12.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2c2633d9ba2e32c85131da23a7ba79aa181a54faeaae9bfbfde8e48fdf39e72b
|
|
| MD5 |
f18b4c6ef5d95ad2674147740c536e48
|
|
| BLAKE2b-256 |
3ef54663c01c555773ca0c9eaa76ada74a1cf4fc6bfc53568fb303f072d88833
|
Provenance
The following attestation bundles were made for parsernaam-0.3.0-py3-none-any.whl:
Publisher:
release.yml on appeler/parsernaam
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
parsernaam-0.3.0-py3-none-any.whl -
Subject digest:
2c2633d9ba2e32c85131da23a7ba79aa181a54faeaae9bfbfde8e48fdf39e72b - Sigstore transparency entry: 2500454408
- Sigstore integration time:
-
Permalink:
appeler/parsernaam@e2c618e18a3ad5ec42d3051bcb209b13094b1a58 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/appeler
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@e2c618e18a3ad5ec42d3051bcb209b13094b1a58 -
Trigger Event:
push
-
Statement type: