The scikit-learn of empirical economics.
Project description
open-econs
Python econometrics with Stata/R parity.
Empirical researchers often use Python for data cleaning, Stata for estimation, R for specialized methods, and LaTeX or custom scripts for output — a fragmented workflow that hurts reproducibility.
open-econs brings familiar empirical economics methods from Stata and R into a unified Python workflow. Every estimator uses a consistent API, with numerical validation against established reference implementations.
- Familiar estimators — OLS, fixed effects, IV/2SLS, logit, probit, difference-in-differences, event studies, RDD, panel models, matching,...
- Tested against Stata/R 217 Stata- and R-parity tests (208 vs Stata, 9 vs R) verify coefficients and
standard errors against reference implementations — at tolerances from
1e-6(typical) up to1e-15for the tightest exact-equivalence cases (IV / Arellano-Bond / synthetic control), mostly1e-6–1e-10. They run in CI on every release, so a numerical-equivalence regression fails the build before it ships. - One consistent interface across all estimators:
.summary(),.tidy(),.vcov(),.predict(),.export(),.to_latex(). - Named pandas outputs — coefficients, standard errors, and diagnostics
return as
pd.Seriesorpd.DataFramewith clear labels - Immutable results — no accidental mutation after estimation
- Reproducible exports — JSON, CSV, LaTeX, HTML from any result
Installation & Quick Start
Requires Python ≥ 3.10.
pip install open-econs # core: OLS, Oaxaca, FE, IV, Logit, Probit
pip install open-econs[plot] # + matplotlib for .plot()
pip install open-econs[nls] # + sympy for nls() (nonlinear least squares)
pip install open-econs[dev,lint] # + development & linting tools
pip install git+https://github.com/qmanhbeo/open-econs.git # latest dev
import open_econs as oe
import pandas as pd
df = pd.DataFrame({
"income": [30, 45, 55, 70, 85, 40, 60, 95],
"education": [10, 12, 14, 16, 18, 11, 15, 20],
"age": [25, 30, 35, 40, 45, 28, 38, 50],
"female": [0, 0, 0, 0, 1, 1, 1, 1],
"province": ["A","A","B","B","C","C","A","B"],
})
r = oe.ols("income ~ education + age", data=df, cluster="province")
print(r.tidy())
# Variable Coef Std Err t P>|t| 0.025 0.975
# 0 Intercept -32.820513 0.730054 -44.956283 0.000000 -34.251392 -31.389633
# 1 education 8.974359 12.843154 0.698766 0.484698 -16.197761 34.146479
# 2 age -1.025641 5.156134 -0.198917 0.842328 -11.131478 9.080196
print(r.summary())
# OLS Regression Results
# ======================================================================
# Dep. Variable: income
# No. Observations: 8
# Df Residuals: 5
# Df Model: 2
# Covariance Type: cluster(province)
# R-squared: 0.991941
# Adj. R-squared: 0.988718
# Condition No.: 3.31e+02
# F-statistic (cluster(province)): 645774.0223
# Prob (F-statistic): 1.548527e-06
# Log-Likelihood: -16.392
# AIC: 38.78
# BIC: 39.02
# ======================================================================
# Variable Coef Std Err t P>|t| 0.025 0.975
# Intercept -32.820513 0.730054 -44.956283 0.000000 -34.251392 -31.389633
# education 8.974359 12.843154 0.698766 0.484698 -16.197761 34.146479
# age -1.025641 5.156134 -0.198917 0.842328 -11.131478 9.080196
# ======================================================================
# Diagnostics:
# Jarque-Bera (chi2=0.498, p=0.7796)
# Breusch-Pagan (LM=5.296, p=0.0708)
# Durbin-Watson: 2.1311
# Ramsey RESET (F=0.140, p=0.8731)
# ======================================================================
Why open-econs?
open-econs keeps everything in one Python environment:
For Stata and R users
If you know the Stata or R command, you already know the open-econs equivalent.
Stata
| Stata | open-econs |
|---|---|
regress |
oe.ols() |
xtreg, fe |
oe.fe() |
ivregress 2sls |
oe.iv() |
logit / probit |
oe.logit() / oe.probit() |
mlogit |
oe.mlogit() (multinomial logit) |
oaxaca |
oe.oaxaca() |
xtabond2 |
oe.abond() |
csdid |
oe.did_cs() |
rdrobust |
oe.rdd() |
teffects psmatch |
oe.psm() |
R
| R | open-econs |
|---|---|
fixest / plm |
oe.fe() / oe.PanelContext() |
AER::ivreg |
oe.iv() |
did |
oe.did_cs() |
MatchIt |
oe.psm() |
See Migrating from Stata for a detailed migration guide, or Migrating from R for the R equivalent.
Step-by-step porting walkthroughs for the causal estimators live in Tutorials:
- RDD (maps to
rdrobust/rddensity) - PSM (maps to
MatchIt/teffects psmatch) - Synthetic Control (maps to
Synth/synth)
Quick Start
import open_econs as oe
import pandas as pd
df = pd.DataFrame({
"income": [30, 45, 55, 70, 85, 40, 60, 95],
"education": [10, 12, 14, 16, 18, 11, 15, 20],
"age": [25, 30, 35, 40, 45, 28, 38, 50],
"female": [0, 0, 0, 0, 1, 1, 1, 1],
"province": ["A","A","B","B","C","C","A","B"],
})
# --- OLS with cluster-robust SEs ---
r = oe.ols("income ~ education + age", data=df, cluster="province")
r.coefficients # pd.Series with named index
r.tidy() # coefficient table as DataFrame
r.predict(df.head(2)) # out-of-sample predictions
print(r.summary()) # printable summary
# --- Logit / Probit ---
r_logit = oe.logit("female ~ education + age", data=df)
r_logit.tidy() # coef, z, P>|z| table
r_logit.margins() # average marginal effects
r_logit.predict(proba=False)
# --- Fixed effects ---
r_fe = oe.fe("income ~ education + age", data=df, entity="province")
r_fe.tidy()
# --- IV / 2SLS ---
r_iv = oe.iv("income ~ education | age", data=df)
r_iv.tidy() # 2SLS coefficients
r_iv.first_stage() # first-stage F-stat
# --- Oaxaca-Blinder decomposition ---
d = oe.oaxaca("income ~ education + age + female", data=df, by="female")
d.explained # covariate-driven gap
d.unexplained # coefficient-driven gap
d.total_gap # female mean - male mean
# --- VIF diagnostics ---
ctx = oe.Context(df)
ctx.vif("income ~ education + age")
# --- Context remembers the dataset ---
ctx.ols("income ~ education + age")
ctx.logit("female ~ education + age")
ctx.probit("female ~ education + age")
# --- Immutability ---
r.f_statistic = 0.0 # AttributeError: OLSResult is immutable
Supported Methods
| Function | Description | Status |
|---|---|---|
ols() / reg() |
OLS with HC1/robust/clustered SEs, multi-way clustering, Newey-West HAC, WLS | Production |
fe() |
Fixed effects (one-way entity, two-way entity + time) | Production |
iv() |
Instrumental variables / 2SLS with first-stage F-stat | Production |
logit() |
Binary logit with .margins(), .predict() |
Production |
probit() |
Binary probit (same API as logit) | Production |
mlogit() |
Multinomial logit (MNLogit) with per-outcome .margins() (dict), .predict(), robust/cluster SEs |
Beta |
oaxaca() |
Oaxaca-Blinder decomposition (two-fold, three-fold; multiple reference types) | Production |
nls() |
Nonlinear least squares (Gauss-Newton via scipy, analytic Jacobian) | Beta |
gmm() |
General linear GMM framework (reuses iv() formula grammar; Hansen J size/power) |
Beta |
abond() |
Arellano-Bond dynamic panel GMM (one/two-step, Windmeijer SEs, collapsed instruments) | Production |
did_cs() |
Callaway-Sant'Anna (2021) staggered DiD (doubly-robust dripw or reg) |
Beta |
did_sa() |
Sun & Abraham (2021) interaction-weighted event-study DiD (cluster-robust VCE) | Beta |
did_gardner() |
Gardner (2022) two-stage DiD (DID2S, cluster-robust IF SEs) | Beta |
rdd() |
Sharp / fuzzy regression discontinuity (local linear, triangular kernel) | Production |
did() / event_study() |
Two-period DiD, event-study with pre-trend diagnostics | Production |
psm() |
Propensity score matching (1:1 nearest-neighbor with replacement, AI 2012 SEs) | Production |
cem() |
Coarsened exact matching (auto or explicit cutpoints, multiple binning methods) | Production |
balance() |
Covariate balance diagnostics (SMD, variance ratio, weighted t-tests) | Production |
rosenbaum_bounds() |
Sensitivity analysis for matched pairs | Beta |
synth() |
Synthetic control (Abadie-Diamond-Hainmueller) core point estimator: nested predictor-weight (V) + donor-weight (W) optimization, gap path, R/Stata parity | Production |
placebo_space() |
Synthetic control placebo-in-space permutation inference (ADH): re-fits synth() once per donor; post/pre MSPE ratio per placebo; permutation p-value; optional exclude_pre_mspe_multiple (space-only) |
Beta |
placebo_time() |
Synthetic control placebo-in-time permutation inference (ADH): re-fits synth() once per candidate pre-treatment date; post/pre MSPE ratio per candidate; permutation p-value |
Beta |
PanelContext(...) |
Pooled/FE/RE/FD/Driscoll-Kraay/Hausman/ABond with remembered entity/time | Production |
Context(...) |
Dataset-scoped workflow with access to all above estimators | Production |
Status reflects validation maturity: Production = machine-precision parity
vs Stata/R on reference fixtures; Beta = validated but with narrower parity
coverage or known edge cases documented in methodology/. Every estimator
ships with parity tests before merge.
Result API
Every estimator returns an object with:
| Method | Returns |
|---|---|
.summary() |
Printable string (also __repr__) |
.tidy() |
pd.DataFrame — coefficient or effect table |
.vcov() |
pd.DataFrame — variance-covariance matrix |
.predict(newdata) |
pd.Series — only on regression models |
.export(path) |
JSON / CSV serialization |
.plot() |
Residual diagnostics plot (requires pip install open-econs[plot]) |
.to_dict() |
dict — full result metadata |
.to_latex() |
LaTeX table string |
.to_html() |
HTML table string |
Design Philosophy
Reproducibility first
Immutable results prevent accidental mutation. All numeric artifacts are named
pd.Series or pd.DataFrame with variable-name indices. Every result can be
exported to JSON, CSV, LaTeX, or HTML.
Familiar econometrics
Methods and defaults follow established practice from Stata and R. If you know
regress, xtreg, or ivregress, you already know how to use the open-econs
equivalent. Default standard errors (HC2) match modern Stata conventions.
One Python workflow
Data processing, estimation, diagnostics, sensitivity analysis, and output generation live in the same environment. No more stitching together separate tools.
Methodology Documentation
Every estimator includes a methodology specification covering:
- Mathematical formulation
- Assumptions
- Estimation procedure
- Inference (all covariance estimators with exact formulas)
- Implementation details
- Stata/R equivalents
- Academic references
See Methodology for the full registry.
Tutorials
Step-by-step walkthroughs for the core estimators:
Comparison
| Library | Strength |
|---|---|
| statsmodels | General statistical models |
| linearmodels | Panel and IV estimation |
| scikit-learn | Machine learning |
| open-econs | Empirical economics workflows with Stata/R parity |
Roadmap
See Roadmap for planned features and development milestones.
Development
pip install -e ".[dev]"
python -m pytest tests/
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file open_econs-1.0.2.tar.gz.
File metadata
- Download URL: open_econs-1.0.2.tar.gz
- Upload date:
- Size: 807.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
26974ddeb466e98d0a829a4b09ed67c2d389accb23d2a52a5c02a6eddb08b0a5
|
|
| MD5 |
1d5c730802427997e39012cffbf24c19
|
|
| BLAKE2b-256 |
db808ff95a5d6a781d08a2d6542944bfd2a9a56e0eb3c6f214bfdd46f93f27f1
|
Provenance
The following attestation bundles were made for open_econs-1.0.2.tar.gz:
Publisher:
publish.yml on qmanhbeo/open-econs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
open_econs-1.0.2.tar.gz -
Subject digest:
26974ddeb466e98d0a829a4b09ed67c2d389accb23d2a52a5c02a6eddb08b0a5 - Sigstore transparency entry: 2170306206
- Sigstore integration time:
-
Permalink:
qmanhbeo/open-econs@2cbede7cb9bc5da0a97a903de5c4e4b461529b1f -
Branch / Tag:
refs/tags/v1.0.2 - Owner: https://github.com/qmanhbeo
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@2cbede7cb9bc5da0a97a903de5c4e4b461529b1f -
Trigger Event:
release
-
Statement type:
File details
Details for the file open_econs-1.0.2-py3-none-any.whl.
File metadata
- Download URL: open_econs-1.0.2-py3-none-any.whl
- Upload date:
- Size: 172.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bc78bb33c45266a37af1e1b38ba1ab1ab3b34ff4bdce814ca9e496d3f324e47b
|
|
| MD5 |
bb738e59da22277f103ab70aa614475d
|
|
| BLAKE2b-256 |
e34f864f701a84812fc930735517f31b26bbbe574c87f9c6c83c7db079a3b65f
|
Provenance
The following attestation bundles were made for open_econs-1.0.2-py3-none-any.whl:
Publisher:
publish.yml on qmanhbeo/open-econs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
open_econs-1.0.2-py3-none-any.whl -
Subject digest:
bc78bb33c45266a37af1e1b38ba1ab1ab3b34ff4bdce814ca9e496d3f324e47b - Sigstore transparency entry: 2170306710
- Sigstore integration time:
-
Permalink:
qmanhbeo/open-econs@2cbede7cb9bc5da0a97a903de5c4e4b461529b1f -
Branch / Tag:
refs/tags/v1.0.2 - Owner: https://github.com/qmanhbeo
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@2cbede7cb9bc5da0a97a903de5c4e4b461529b1f -
Trigger Event:
release
-
Statement type: