This release is a pre-release and may not be stable for production use.
pnadtables
Fairness benchmark tasks on Brazil's PNAD Contínua (IBGE), with an API inspired by
folktables. Microdata are accessed
directly from the official IBGE FTP through pnadium; Base dos Dados / BigQuery is an
optional backend.
Leia em português.
Status — pre-release, partially validated against real data
This package is not ready for use in published results. Version 0.1.0.dev1 is a
working pre-release: its logic is tested offline, and the SP/2024Q1-Q2 and MT/2024Q1
official files passed real-data validation (SP Q1,
SP Q2, MT Q1). The
experimental baseline now covers a household-grouped split
(report), temporal transfer
(report), and geographic transfer
(report). The primary baseline also
has conditional 95% intervals from all 200 PNADC bootstrap weights
(report). Training uncertainty and
reproduction by a second person remain open.
No release has been tagged and no DOI exists.
For a detailed Portuguese account of the project history, architecture, data and
experiments, see the complete project report.
Open gates before v0.1.0 (see pnadtables-plano.md):
| Gate | What it requires | Status |
|---|---|---|
| G0 | Git repository, correct metadata, no secrets | ✅ |
| G1 | Schema and value domains confirmed against the IBGE dictionary and real data | ✅ SP/2024Q1-Q2 and MT/2024Q1 passed |
| G2 | Deterministic encoding, missingness policy, weight validation, integration tests | 🟨 local validation implemented; real-data integration still open |
| G3 | Leakage-free splits, survey-design protocol, metrics with uncertainty | 🟨 splits and conditional bootstrap intervals done; training uncertainty remains open |
| G4 | Reproducible real-data baseline with a results table | 🟨 real-data reports produced; second-person reproduction remains open |
| G5 | CI green on supported Python versions, clean install from artifact | ⬜ |
| G6 | Datasheet, limitations, provenance, verified licences/terms | ⬜ |
| G7 | Public repo, v0.1.0 tag, archived release, DOI |
⬜ |
| G8 | PyPI publication, verified pip install pnadtables |
⬜ |
from pnadtables import PNADCDataSource, PNADEmployment, audit_report
src = PNADCDataSource(
ano=2024, trimestre=1, ufs=["MT", "SP"],
caminho="dados", salvar=True, formato="csv",
# incluir_pesos_replicados=True adds the 200 bootstrap weights for CIs.
)
X, y, group, weight = PNADEmployment.df_to_pandas(src.get_data())
# 4-tuple: folktables returns (X, y, group); pnadtables adds the PNADC survey weight.
# Code written against folktables will NOT unpack this without modification.
# Nominal features are pandas categoricals; fit an encoder on training data only.
df_to_numpy() refuses nominal features instead of inventing integer codes. Use
df_to_pandas() with a training-only preprocessing pipeline (see example_audit.py).
Why
Most empirical fairness research is calibrated on U.S. Census data (UCI Adult, then
folktables). Whether its conclusions transfer to the Global South is, as far as we know,
largely untested — but this is a hypothesis we have not yet verified with a systematic
literature search (ACM DL, IEEE Xplore, Scopus/OpenAlex, arXiv, Google Scholar). Do not
cite this README as evidence that no comparable benchmark exists. pnadtables adapts the
BasicProblem design to Brazil's official labour survey and targets two things:
- Racial taxonomies are not translatable 1:1. IBGE uses five self-declared categories (Branca, Preta, Amarela, Parda, Indígena), plus code 9 for "Ignorada" (non-response). Collapsing Preta + Parda into a binary Black/white axis is a theoretical choice, not a data fact. The package ships both encodings so results can be reported under each.
- Complex survey design. PNADC is a weighted, stratified, clustered sample.
weightalone yields weighted point estimates. Withincluir_pesos_replicados=True, the baseline uses all 200 bootstrap weights for conditional intervals under the sampling design. These intervals keep predictions fixed and do not cover training uncertainty.
Not a drop-in folktables replacement
Name parity is not construct parity. ACS and PNADC differ in universe, periodicity, variables and sample design, and the race/colour categories are not semantically interchangeable. A Brazil–U.S. comparison needs an explicit harmonisation table (construct, universe, income window, features, categories, geography, design, metric), not just running tasks with similar names.
Tasks
| Task | Conceptual reference | Universe | Target |
|---|---|---|---|
PNADEmployment |
ACSEmployment | age ≥ 14 | employed (vd4002 == 1) |
PNADIncome |
ACSIncome | employed, age ≥ 14, income > 0 | monthly income > threshold (parameter) |
PNADIncomeRegression |
continuous complement | employed, age ≥ 14, income > 0 | monthly income in BRL |
Features: age, sex, race/colour, education level, state (UF). Protected attribute:
race/colour. Build new tasks with BasicProblem(features, target, target_transform, group, ...).
Install
pip install git+https://github.com/Brunooliveirab/pnadtables
pip install "pnadtables[baseline]" # adds scikit-learn for pnadtables-baseline
pip install "pnadtables[examples]" # adds scikit-learn + shap for example_audit.py
The default backend downloads the public quarterly file from IBGE through pnadium. It
requires no account, credentials or billing project. Selecting fewer variables reduces
memory use, although the compressed quarterly archive still has to be downloaded.
from pnadtables import PNADCDataSource
src = PNADCDataSource(ano=2024, trimestre=1, ufs=None) # all 27 states
df = src.get_data()
Run the reproducible baseline with:
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet
# Continuous income: log1p linear fit, evaluated in BRL with MAE/RMSE.
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet \
--tarefa income_regression
# Train on SP/2024Q1 and evaluate on SP/2024Q2, removing repeated households.
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet \
--teste-externo dados/pnadtables_pnadc_2024T2_SP.parquet \
--saida docs/baseline-temporal-sp-2024t1-t2.md
The table below is the historical binary-race result from before 0.1.0.dev1; reproduce
it with --race-mode binary. The current default uses five groups.
| Model | Weighted accuracy (95% CI) | Weighted ROC AUC (95% CI) |
|---|---|---|
| Weighted majority | 0.617 [0.604, 0.630] | 0.500 [0.500, 0.500] |
| Logistic regression with protected attribute | 0.690 [0.678, 0.702] | 0.721 [0.708, 0.734] |
| Logistic regression without protected attribute | 0.690 [0.678, 0.701] | 0.716 [0.703, 0.730] |
These SP/2024Q1 intervals use the 200 bootstrap weights and condition on the fitted predictions; they do not include training uncertainty. See the methodology.
Optional BigQuery backend
Install pnadtables[bigquery] and use the explicit backend when remote columnar querying
is preferable. This path requires a billed Google Cloud project. It performs a free dry
run and enforces maximum_bytes_billed before executing the query.
from pnadtables import PNADCBigQueryDataSource
src = PNADCBigQueryDataSource("my-gcp-project", 2024, 1, ufs=["MT", "SP"])
print(src.dry_run()) # free: what this query would cost, before running it
df = src.get_data() # refuses above DEFAULT_MAX_GB (5 GB)
df = src.get_data(max_gb=20) # raise the ceiling deliberately
df = src.get_data(max_gb=None) # disable both guards — know what you are doing
See BigQuery pricing and
cost controls. Never
replace the selected column list with SELECT *.
Before you publish numbers
- Run
pnadtables-validar-dicionario --snapshot docs/pnadc-dicionario.snapshot.jsonfor the versioned evidence, or use--escreverto refresh it from IBGE. Validate a downloaded quarter's domains before publishing. If using the optional BigQuery backend, also runinspect_bigquery_schema(billing_project_id)to verify that ingestion. Runpnadtables-validar-dados <arquivo.parquet>on every experimental vintage. - Set the income threshold in
make_pnad_income(threshold=...)to the reference year of your data. - Use the five self-declared race/colour categories as the primary analysis (the API and
CLI default), then run
--race-mode binaryas a White vs Black (Black + Brown) sensitivity analysis and quantify exclusions. - Report both weighted and unweighted metrics as a sensitivity analysis. Neither is universally the correct headline number — say which question each one answers.
- Do not use a naive random split: people in the same household are correlated and PNADC
rotates households. Use the household-grouped baseline or
--teste-externofor a temporal/geographic evaluation; repeatedcod_famvalues are removed from the test set.
See DATASHEET.md (Gebru et al. structure) for the full methodological record.
Metrics
selection_rates, disparate_impact, rate_by_group (TPR/FPR), audit_report — all with
optional survey weights. Inputs are validated for length, missing/non-binary values and
invalid weights; the report includes raw counts, weight totals, rate differences and
ratios. By default disparate_impact uses the highest-rate group as
reference, so ratios fall in (0, 1] and the 0.8 threshold is the relevant one; the
symmetric 1.25 threshold only applies when you pass a fixed reference yourself. The 4/5
rule is a contextual screening heuristic, not an automatic legal diagnosis of
discrimination.
For the continuous task, regression_report reports actual/predicted means, signed error,
MAE and RMSE by group. These are descriptive diagnostics, not causal conclusions. Removing
race/colour from the feature matrix does not remove proxy information carried by geography,
education or other correlated variables.
The baseline is also a direct Python API:
from pnadtables import executar_baseline
result = executar_baseline(df, tarefa="income_regression", race_mode="five")
example_audit.py runs the full loop: data → model → weighted audit → SHAP attributions.
SHAP is illustrative only; a standalone feature-importance ranking does not support claims
about a "right to explanation".
Intended use
Research and auditing of algorithmic systems. Not for decisions about individuals, and not a substitute for substantive research on inequality. Do not attempt to re-identify respondents.
Development
pip install -e ".[dev]"
pytest
Tests run offline with synthetic PNADC-shaped data and a mocked IBGE download. Passing them is not evidence that every real-data vintage has the expected domains (gate G1).
Citation
See CITATION.cff. Note that no version has been released yet. Please also cite folktables:
Ding, F., Hardt, M., Miller, J., & Schmidt, L. (2021). Retiring Adult: New Datasets for Fair Machine Learning. NeurIPS 34.
License
MIT for this package's code. Microdata © IBGE; this package redistributes no data. The optional backend accesses the Base dos Dados copy. Attribution and applicable access terms must be verified before publication (gate G6).
Release files for pnadtables 0.1.0.dev1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pnadtables-0.1.0.dev1.tar.gz | 96.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pnadtables-0.1.0.dev1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 139.1 kB
Release files / pnadtables-0.1.0.dev1.tar.gz
| Download URL | pnadtables-0.1.0.dev1.tar.gz |
|---|---|
| Size | 96.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5b9ad8cbf07929431054396a6b5c942a35845b80cdc2b13db8d2443918baf046
|
|
BLAKE2b-256 checksum How to use checksums |
d26e5bdaf8e7d33c65a7f9a01bcea2fb8e4a8383cd10220cf31e0d55a84863f8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.7
|
Release files / pnadtables-0.1.0.dev1-py3-none-any.whl
| Download URL | pnadtables-0.1.0.dev1-py3-none-any.whl |
|---|---|
| Size | 42.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
96fd478bbb7b5d7326298ace88a3348d3d408c76f1de90c7241399727823c414
|
|
BLAKE2b-256 checksum How to use checksums |
31366474bacf81b3459afb1ac02ba3c4faab09b28e883b9354ae38a3eae48191
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.7
|