Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

pnadtables

CI License: MIT

Fairness benchmark tasks on Brazil's PNAD Contínua (IBGE), with an API inspired by folktables. Microdata are accessed directly from the official IBGE FTP through pnadium; Base dos Dados / BigQuery is an optional backend.

Leia em português.

Status — pre-release, partially validated against real data

This package is not ready for use in published results. Version 0.1.0.dev1 is a working pre-release: its logic is tested offline, and the SP/2024Q1-Q2 and MT/2024Q1 official files passed real-data validation (SP Q1, SP Q2, MT Q1). The experimental baseline now covers a household-grouped split (report), temporal transfer (report), and geographic transfer (report). The primary baseline also has conditional 95% intervals from all 200 PNADC bootstrap weights (report). Training uncertainty and reproduction by a second person remain open. No release has been tagged and no DOI exists. For a detailed Portuguese account of the project history, architecture, data and experiments, see the complete project report.

Open gates before v0.1.0 (see pnadtables-plano.md):

Gate What it requires Status
G0 Git repository, correct metadata, no secrets ✅
G1 Schema and value domains confirmed against the IBGE dictionary and real data ✅ SP/2024Q1-Q2 and MT/2024Q1 passed
G2 Deterministic encoding, missingness policy, weight validation, integration tests 🟨 local validation implemented; real-data integration still open
G3 Leakage-free splits, survey-design protocol, metrics with uncertainty 🟨 splits and conditional bootstrap intervals done; training uncertainty remains open
G4 Reproducible real-data baseline with a results table 🟨 real-data reports produced; second-person reproduction remains open
G5 CI green on supported Python versions, clean install from artifact ⬜
G6 Datasheet, limitations, provenance, verified licences/terms ⬜
G7 Public repo, v0.1.0 tag, archived release, DOI ⬜
G8 PyPI publication, verified pip install pnadtables ⬜
from pnadtables import PNADCDataSource, PNADEmployment, audit_report

src = PNADCDataSource(
    ano=2024, trimestre=1, ufs=["MT", "SP"],
    caminho="dados", salvar=True, formato="csv",
    # incluir_pesos_replicados=True adds the 200 bootstrap weights for CIs.
)
X, y, group, weight = PNADEmployment.df_to_pandas(src.get_data())
# 4-tuple: folktables returns (X, y, group); pnadtables adds the PNADC survey weight.
# Code written against folktables will NOT unpack this without modification.
# Nominal features are pandas categoricals; fit an encoder on training data only.

df_to_numpy() refuses nominal features instead of inventing integer codes. Use df_to_pandas() with a training-only preprocessing pipeline (see example_audit.py).

Why

Most empirical fairness research is calibrated on U.S. Census data (UCI Adult, then folktables). Whether its conclusions transfer to the Global South is, as far as we know, largely untested — but this is a hypothesis we have not yet verified with a systematic literature search (ACM DL, IEEE Xplore, Scopus/OpenAlex, arXiv, Google Scholar). Do not cite this README as evidence that no comparable benchmark exists. pnadtables adapts the BasicProblem design to Brazil's official labour survey and targets two things:

  • Racial taxonomies are not translatable 1:1. IBGE uses five self-declared categories (Branca, Preta, Amarela, Parda, Indígena), plus code 9 for "Ignorada" (non-response). Collapsing Preta + Parda into a binary Black/white axis is a theoretical choice, not a data fact. The package ships both encodings so results can be reported under each.
  • Complex survey design. PNADC is a weighted, stratified, clustered sample. weight alone yields weighted point estimates. With incluir_pesos_replicados=True, the baseline uses all 200 bootstrap weights for conditional intervals under the sampling design. These intervals keep predictions fixed and do not cover training uncertainty.

Not a drop-in folktables replacement

Name parity is not construct parity. ACS and PNADC differ in universe, periodicity, variables and sample design, and the race/colour categories are not semantically interchangeable. A Brazil–U.S. comparison needs an explicit harmonisation table (construct, universe, income window, features, categories, geography, design, metric), not just running tasks with similar names.

Tasks

Task Conceptual reference Universe Target
PNADEmployment ACSEmployment age ≥ 14 employed (vd4002 == 1)
PNADIncome ACSIncome employed, age ≥ 14, income > 0 monthly income > threshold (parameter)
PNADIncomeRegression continuous complement employed, age ≥ 14, income > 0 monthly income in BRL

Features: age, sex, race/colour, education level, state (UF). Protected attribute: race/colour. Build new tasks with BasicProblem(features, target, target_transform, group, ...).

Install

pip install git+https://github.com/Brunooliveirab/pnadtables
pip install "pnadtables[baseline]"   # adds scikit-learn for pnadtables-baseline
pip install "pnadtables[examples]"   # adds scikit-learn + shap for example_audit.py

The default backend downloads the public quarterly file from IBGE through pnadium. It requires no account, credentials or billing project. Selecting fewer variables reduces memory use, although the compressed quarterly archive still has to be downloaded.

from pnadtables import PNADCDataSource

src = PNADCDataSource(ano=2024, trimestre=1, ufs=None)  # all 27 states
df = src.get_data()

Run the reproducible baseline with:

pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet

# Continuous income: log1p linear fit, evaluated in BRL with MAE/RMSE.
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet \
  --tarefa income_regression

# Train on SP/2024Q1 and evaluate on SP/2024Q2, removing repeated households.
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet \
  --teste-externo dados/pnadtables_pnadc_2024T2_SP.parquet \
  --saida docs/baseline-temporal-sp-2024t1-t2.md

The table below is the historical binary-race result from before 0.1.0.dev1; reproduce it with --race-mode binary. The current default uses five groups.

Model Weighted accuracy (95% CI) Weighted ROC AUC (95% CI)
Weighted majority 0.617 [0.604, 0.630] 0.500 [0.500, 0.500]
Logistic regression with protected attribute 0.690 [0.678, 0.702] 0.721 [0.708, 0.734]
Logistic regression without protected attribute 0.690 [0.678, 0.701] 0.716 [0.703, 0.730]

These SP/2024Q1 intervals use the 200 bootstrap weights and condition on the fitted predictions; they do not include training uncertainty. See the methodology.

Optional BigQuery backend

Install pnadtables[bigquery] and use the explicit backend when remote columnar querying is preferable. This path requires a billed Google Cloud project. It performs a free dry run and enforces maximum_bytes_billed before executing the query.

from pnadtables import PNADCBigQueryDataSource

src = PNADCBigQueryDataSource("my-gcp-project", 2024, 1, ufs=["MT", "SP"])

print(src.dry_run())     # free: what this query would cost, before running it
df = src.get_data()      # refuses above DEFAULT_MAX_GB (5 GB)
df = src.get_data(max_gb=20)   # raise the ceiling deliberately
df = src.get_data(max_gb=None) # disable both guards — know what you are doing

See BigQuery pricing and cost controls. Never replace the selected column list with SELECT *.

Before you publish numbers

  1. Run pnadtables-validar-dicionario --snapshot docs/pnadc-dicionario.snapshot.json for the versioned evidence, or use --escrever to refresh it from IBGE. Validate a downloaded quarter's domains before publishing. If using the optional BigQuery backend, also run inspect_bigquery_schema(billing_project_id) to verify that ingestion. Run pnadtables-validar-dados <arquivo.parquet> on every experimental vintage.
  2. Set the income threshold in make_pnad_income(threshold=...) to the reference year of your data.
  3. Use the five self-declared race/colour categories as the primary analysis (the API and CLI default), then run --race-mode binary as a White vs Black (Black + Brown) sensitivity analysis and quantify exclusions.
  4. Report both weighted and unweighted metrics as a sensitivity analysis. Neither is universally the correct headline number — say which question each one answers.
  5. Do not use a naive random split: people in the same household are correlated and PNADC rotates households. Use the household-grouped baseline or --teste-externo for a temporal/geographic evaluation; repeated cod_fam values are removed from the test set.

See DATASHEET.md (Gebru et al. structure) for the full methodological record.

Metrics

selection_rates, disparate_impact, rate_by_group (TPR/FPR), audit_report — all with optional survey weights. Inputs are validated for length, missing/non-binary values and invalid weights; the report includes raw counts, weight totals, rate differences and ratios. By default disparate_impact uses the highest-rate group as reference, so ratios fall in (0, 1] and the 0.8 threshold is the relevant one; the symmetric 1.25 threshold only applies when you pass a fixed reference yourself. The 4/5 rule is a contextual screening heuristic, not an automatic legal diagnosis of discrimination.

For the continuous task, regression_report reports actual/predicted means, signed error, MAE and RMSE by group. These are descriptive diagnostics, not causal conclusions. Removing race/colour from the feature matrix does not remove proxy information carried by geography, education or other correlated variables.

The baseline is also a direct Python API:

from pnadtables import executar_baseline

result = executar_baseline(df, tarefa="income_regression", race_mode="five")

example_audit.py runs the full loop: data → model → weighted audit → SHAP attributions. SHAP is illustrative only; a standalone feature-importance ranking does not support claims about a "right to explanation".

Intended use

Research and auditing of algorithmic systems. Not for decisions about individuals, and not a substitute for substantive research on inequality. Do not attempt to re-identify respondents.

Development

pip install -e ".[dev]"
pytest

Tests run offline with synthetic PNADC-shaped data and a mocked IBGE download. Passing them is not evidence that every real-data vintage has the expected domains (gate G1).

Citation

See CITATION.cff. Note that no version has been released yet. Please also cite folktables:

Ding, F., Hardt, M., Miller, J., & Schmidt, L. (2021). Retiring Adult: New Datasets for Fair Machine Learning. NeurIPS 34.

License

MIT for this package's code. Microdata © IBGE; this package redistributes no data. The optional backend accesses the Base dos Dados copy. Attribution and applicable access terms must be verified before publication (gate G6).

Release files for pnadtables 0.1.0.dev1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pnadtables 0.1.0.dev1
File Size Uploaded
pnadtables-0.1.0.dev1.tar.gz 96.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pnadtables 0.1.0.dev1
File Interpreter ABI Platform
pnadtables-0.1.0.dev1-py3-none-any.whl Python 3 none any Details

Total release size: 139.1 kB

Release files / pnadtables-0.1.0.dev1.tar.gz

Download URL pnadtables-0.1.0.dev1.tar.gz
Size 96.7 kB
Tags Source
SHA-256 checksum
How to use checksums
5b9ad8cbf07929431054396a6b5c942a35845b80cdc2b13db8d2443918baf046
BLAKE2b-256 checksum
How to use checksums
d26e5bdaf8e7d33c65a7f9a01bcea2fb8e4a8383cd10220cf31e0d55a84863f8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.7

Release files / pnadtables-0.1.0.dev1-py3-none-any.whl

Download URL pnadtables-0.1.0.dev1-py3-none-any.whl
Size 42.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
96fd478bbb7b5d7326298ace88a3348d3d408c76f1de90c7241399727823c414
BLAKE2b-256 checksum
How to use checksums
31366474bacf81b3459afb1ac02ba3c4faab09b28e883b9354ae38a3eae48191
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.7

Release history Release notifications | RSS feed

This release

0.1.0.dev1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page