geoestimate
Design-aware estimates for equal-probability observation samples.
geoestimate estimates means, population totals, ratios of totals, and means
of row-level ratios. It reports iid, cluster-sandwich, or pairs-bootstrap
uncertainty. The stable API assumes that rows have equal inclusion
probabilities.
The package does not implement unequal-probability weights, PPS, GRTS, stratification, finite population corrections, nonresponse adjustments, annotation-subsampling corrections, or walk-spacing corrections.
Install
pip install geoestimate
Estimate from a sample
import pandas as pd
from geoestimate import Sample
frames = pd.DataFrame(
{
"n_women": [3, 4, 2, 5],
"n_people": [10, 10, 10, 10],
"itinerary_id": [0, 0, 1, 1],
}
)
sample = Sample(frames, cluster="itinerary_id")
mean_people = sample.mean("n_people")
total_people = sample.total("n_people", population_size=50_000)
people_share = sample.ratio("n_women", "n_people")
location_share = sample.mean_of_ratios("n_women", "n_people")
print(people_share.summary())
Declare a cluster only when it identifies independent sampling or collection units. A declared cluster makes cluster-sandwich inference with a Student-t interval the default. Without a cluster, the default is an iid standard error with a normal interval.
Choose the estimand
mean("n_people") estimates the average count per population unit. For a
binary variable, the mean is a population proportion.
total("n_people", population_size=N) estimates N * mean(n_people). The
population size counts the same row-level units represented by the sample. The
method does not apply a finite population correction.
ratio("n_women", "n_people") estimates:
sum(n_women) / sum(n_people)
This ratio weights rows by their denominator. Individual denominators may be zero, but they must be nonnegative and their sample total must be positive.
mean_of_ratios("n_women", "n_people") estimates:
mean(n_women / n_people)
This estimand gives every row equal weight. Every denominator must be positive.
Filter the DataFrame before constructing Sample when the target population
excludes rows with zero denominators.
Select inference
The default inference="design" follows the declared sample design. You can
request a method explicitly:
bootstrap = sample.ratio(
"n_women",
"n_people",
inference="bootstrap",
bootstrap_reps=2_000,
seed=42,
)
The bootstrap resamples declared clusters, or individual rows when no cluster
is declared. It reports the standard deviation of the bootstrap estimates and
a percentile interval. Explicit inference="iid" is available as a sensitivity
comparison for clustered samples.
Read files and use the command line
Sample.from_file() accepts Parquet, CSV, compressed CSV, and TSV files:
sample = Sample.from_file("frames.parquet", cluster="itinerary_id")
result = sample.ratio("n_women", "n_people")
The command line exposes the same four estimands:
geoestimate mean frames.parquet --variable n_people --cluster itinerary_id
geoestimate total frames.parquet --variable n_people --population-size 50000
geoestimate ratio frames.parquet --numerator n_women --denominator n_people
geoestimate mean-of-ratios frames.parquet --numerator n_women --denominator n_people
Experimental validation tools
geoestimate.spatial, geoestimate.simulate, and geoestimate.pipeline help
diagnose dependence and validate collection designs. Their interfaces may
change before the stable inference API does.
Install the optional pipeline dependencies to validate a real geo-sampling
to allocator geometry:
pip install "geoestimate[pipeline]"
python examples/validate_with_allocator.py
License
MIT
Metadata
Release files for geoestimate 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| geoestimate-0.1.1.tar.gz | 29.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| geoestimate-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 62.8 kB
Release files / geoestimate-0.1.1.tar.gz
| Download URL | geoestimate-0.1.1.tar.gz |
|---|---|
| Size | 29.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ef2e62b1f6a2841d80cac0ea6393dd6607470859a2e99cdb9ee7f4d14194d657
|
|
BLAKE2b-256 checksum How to use checksums |
6e41e9c57117ba31794ddd9b7b0b944c667c798c305de73a920b203808bf080b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / geoestimate-0.1.1-py3-none-any.whl
| Download URL | geoestimate-0.1.1-py3-none-any.whl |
|---|---|
| Size | 33.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
009ee6f933643fd80b2def9377f6889de578bb785cad00ad9f1d06486b1ae166
|
|
BLAKE2b-256 checksum How to use checksums |
f651157deb42688e60e42ad09d2a8e6aeb62a707ed6c1367222ca766d3098ccf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log