Skip to main content

indian-address-parser

Parse raw, unstructured Indian address strings into 13 structured fields. Three model backends, all downloaded automatically from Hugging Face — this package ships only inference code, no weights:

  • tinybert (default) — a full fine-tune of huawei-noah/TinyBERT_General_4L_312D (model card). ~14M params, token classification (BIO tagging) rather than JSON generation — by far the smallest and fastest option.
  • t5 — a full fine-tune of google/flan-t5-small (model card). ~77M params, single download, a couple points more accurate than tinybert, at noticeably higher inference cost.
  • qwen — a LoRA adapter on Qwen/Qwen3-0.6B (model card). ~596M params, needs both the adapter and the base model, the most accurate option.
Input:  "FLAT NO.32, UTTARA TOWERS, MG ROAD GUWAHATI , Kamrup Unclassified AS 781029"
Output: {"houseNumber": "FLAT NO.32", "houseName": "UTTARA TOWERS", "poi": null,
         "street": "MG ROAD", "subsubLocality": null, "subLocality": null, "locality": null,
         "village": null, "subDistrict": null, "district": "Kamrup", "city": "GUWAHATI",
         "state": "AS", "pincode": "781029"}

Install

pip install indian-address-parser

Usage

Python

from indian_address_parser import AddressParser

parser = AddressParser()  # default backend "tinybert" — downloads model weights from HF on first use
result = parser.parse("FLAT NO.32, UTTARA TOWERS, MG ROAD GUWAHATI , Kamrup Unclassified AS 781029")
print(result)

# Batch
results = parser.parse_batch([addr1, addr2, addr3])

# A couple points more accurate, slower
parser = AddressParser(backend="t5")

# Most accurate model
parser = AddressParser(backend="qwen")

CLI

# Single address (defaults to the tinybert backend)
indian-address-parser "FLAT NO.32, UTTARA TOWERS, MG ROAD GUWAHATI , Kamrup Unclassified AS 781029"

# Batch from stdin
cat addresses.txt | indian-address-parser --stdin

# Batch from a file, JSONL output
indian-address-parser --file addresses.txt --out results.jsonl

# A couple points more accurate, slower
indian-address-parser --backend t5 "..."

# Most accurate model
indian-address-parser --backend qwen "..."

Fields

houseNumber, houseName, poi, street, subsubLocality, subLocality,
locality, village, subDistrict, district, city, state, pincode

Any field not present in the address is null. For the generative backends (t5, qwen), output that can't be parsed as JSON leaves all fields null and a _parse_error key holds the raw model output; tinybert is token classification, not generation, so it always produces a well-formed dict and _parse_error never applies.

Choosing a backend

tinybert (default) t5 qwen
Base model TinyBERT_General_4L_312D google/flan-t5-small Qwen/Qwen3-0.6B + LoRA
Task token classification (BIO) JSON generation JSON generation
Params ~14M ~77M ~596M
Download single model single model adapter + base model
Mean field accuracy (held-out test) 78.8% 80.6% 82.4%

tinybert is the default — at ~14M params (5x smaller than t5, 40x smaller than qwen) it's by far the cheapest to download and run, a single forward pass instead of autoregressive generation, and still lands within 2-4 points of both larger models on mean field accuracy (though its village/subLocality recall is essentially 0% — see its model card for the full breakdown). Reach for t5 when those couple of points matter and the extra inference cost is acceptable, or qwen when accuracy matters most (its edge is largest on the lower-recall fields like poi, subDistrict, subLocality).

Model details, evaluation metrics, and known limitations

See the model cards for training data, config, full per-field evaluation results, and known limitations (locality/subLocality/subsubLocality/village field-boundary ambiguity, etc.):

Datasets

Comparison to other models

benchmarks/ has two head-to-head comparisons against Shiprocket's models on a shared, 237-example held-out gold test set (both models use different field taxonomies than ours, so only overlapping fields are scored):

Apple Silicon (MLX) users

This package uses transformers (+peft for the qwen backend), which works on CUDA, MPS, and CPU but is not the fastest path on Apple Silicon. For MLX-native inference of the qwen backend, see the mlx/ subfolder of its Hugging Face repo instead.

License

Apache 2.0 (matching the base models: google/flan-t5-small, Qwen/Qwen3-0.6B, huawei-noah/TinyBERT_General_4L_312D).

Metadata

Release files for indian-address-parser 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for indian-address-parser 0.3.0
File Size Uploaded
indian_address_parser-0.3.0.tar.gz 98.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for indian-address-parser 0.3.0
File Interpreter ABI Platform
indian_address_parser-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 112.7 kB

Release files / indian_address_parser-0.3.0.tar.gz

Download URL indian_address_parser-0.3.0.tar.gz
Size 98.8 kB
Tags Source
SHA-256 checksum
How to use checksums
1f876a1728e9426876648c7008cf1b569c08269f3c00553f4e2a744ca9f63398
BLAKE2b-256 checksum
How to use checksums
054a8af1c27cea159d9e9f700aa9e06869ef8199ca9b0adc4ba439f21a2cb9e9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release files / indian_address_parser-0.3.0-py3-none-any.whl

Download URL indian_address_parser-0.3.0-py3-none-any.whl
Size 14.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fc297428f74e342d28affab2ca7dd1185f1d8aec28389c6ac8fc923565c76f6c
BLAKE2b-256 checksum
How to use checksums
1425b6e9fb4d851a9578cc392add475d8a90cfa7866da04e79647a7356139d9a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page