microdf
Weighted pandas DataFrames and Series for survey microdata analysis.
Why this exists
Survey microdata comes with weights, and analysing it in pandas means getting two things right that pandas will not do for you.
The estimators are not the obvious ones. A weighted median is not the median of
weighted values. Weighted variance requires choosing between treating weights as
frequencies or as precision. A top-1% share requires deciding what happens to a
record that straddles the cutoff. Each of these is a decision, and hand-rolling it
per analysis means making it differently each time. microdf makes each choice
once, documents it, and tests it — quantiles follow the inverse CDF so they can be
checked against R's survey::svyquantile, and variance treats weights as
frequencies so integer weights agree with numpy on the replicated sample.
Weights have to survive the pipeline. Before any estimator runs, weights must
stay aligned with their rows through merges, filters, grouping and reindexing. When
they do not, nothing raises. The pipeline completes and returns a plausible wrong
number. This is the harder of the two problems, and it is why microdf carries
weights inside the object rather than beside it.
If you are computing a poverty rate or a Gini on weighted survey data, those are the two ways to get a believable-looking wrong answer.
Key Features
- MicroDataFrame: A pandas DataFrame with an integrated weight column
- MicroSeries: A pandas Series with integrated weights
- Weighted operations: All aggregations (sum, mean, median, etc.) automatically use weights
- Inequality metrics: Built-in Gini coefficient calculation
- Poverty analysis: Integrated poverty rate and gap calculations
Installation
Install with:
pip install microdf-python
Or for development:
pip install git+https://github.com/PolicyEngine/microdf.git
Usage
import microdf as mdf
import pandas as pd
# Create sample data with weights
df = pd.DataFrame(
{"income": [10_000, 20_000, 30_000, 40_000, 50_000], "weights": [1, 2, 3, 2, 1]}
)
# Create a MicroDataFrame
mdf_df = mdf.MicroDataFrame(df, weights="weights")
# All operations are weight-aware
print(mdf_df.income.mean()) # Weighted mean
print(mdf_df.income.gini()) # Gini coefficient
Questions
Contact the maintainer, Max Ghenis (max@policyengine.org).
Citation
You may cite the source of your analysis as "microdf release #.#.#, author's calculations."
Metadata
Release files for microdf-python 1.5.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| microdf_python-1.5.6.tar.gz | 61.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| microdf_python-1.5.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 130.3 kB
Release files / microdf_python-1.5.6.tar.gz
| Download URL | microdf_python-1.5.6.tar.gz |
|---|---|
| Size | 61.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ffd970002b94a7161cc8a98d0b0e51926008781b8207bed7349fd715f6bcf816
|
|
BLAKE2b-256 checksum How to use checksums |
907447d07f5ebc8b9bfe18d8e122981eb3d91cc9e2d36f01cf55693701cef2cf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / microdf_python-1.5.6-py3-none-any.whl
| Download URL | microdf_python-1.5.6-py3-none-any.whl |
|---|---|
| Size | 68.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e53fe722a0fa18d72d22f1e16213f5d4926b234f4492faa140e3ef88ade05fe6
|
|
BLAKE2b-256 checksum How to use checksums |
232421a140ef9ab1e96bfc8beda674bfcf0dd32375deb40eb7673c5f9c813c40
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|