Skip to main content

causalrl

CI Docs PyPI License: MIT Python 3.11+ Ruff

Causal reinforcement learning: the 9-task causal RL taxonomy, made runnable.

causalrl supplies the causal layer for sequential decisions, not a new trainer: agents that plan inside a given or learned structural causal model, graph algorithms for causal bandits, demonstration environments, and explicit-latent SCMs with see (L1), do (L2), and counterfactual (L3) queries, organised around the 9-task taxonomy of causal RL. Train a policy however you like, then hand its actions to certify_policy: it bounds whether the value improvement over the logging policy survives hidden confounding, and abstains when it cannot.

Scope is explicit and enforced in code. Out-of-class identification queries raise NotIdentifiableError with the witnessing hedge (or return None for the conservative helpers) rather than guessing a formula. The two halves also sit at deliberately different maturities: the planners and environments are demo-scale — tabular to modest function approximation, on synthetic worlds built to isolate one failure mode each — while the decision, certificate and off-policy-evaluation layers run on real data. The examples/causal_mbrl_*.py scripts are the evidence: on NHEFS, LaLonde and Twins they fit an agent, call .act() per unit and certify the resulting policy; on the Open Bandit Dataset and Coat they run the off-policy-evaluation and sensitivity kernels (certify_policy, msm_policy_value_bounds) against measured ground truth. See Guarantees & Scope.

Install

pip install causalrl            # core: graph, POMIS, tabular agents/environments
pip install "causalrl[torch]"   # + SCM sampling, neural mechanisms, Torch-backed demos

From a clone, for development:

uv sync --extra dev             # tests, lint, typing, notebooks
uv sync --extra docs            # local documentation site and API reference

The core graph, POMIS, tabular-agent, and tabular-environment surfaces do not require PyTorch; SCM sampling, neural mechanisms, and structural-bandit environments do. Full documentation: https://raphaelrrcoelho.github.io/causalrl/.

Quickstart

A causal agent that conditions on its "intuition" beats a confounding-naive agent on the Multi-Armed Bandit with Unobserved Confounders — even though both arms have identical interventional means.

from causalrl.agents.bandits import CausalThompsonSampling
from causalrl.envs.suite.mabuc import MABUCEnv

env = MABUCEnv(seed=1)
agent = CausalThompsonSampling(n_arms=2, n_contexts=2, seed=0)

obs, _ = env.reset(seed=1)
for _ in range(8000):
    action = agent.act(obs)
    _, reward, _, _, _ = env.step(action)
    agent.update(obs, action, reward)
    obs, _ = env.reset()
# CausalThompsonSampling -> ~0.75 reward/step; any confounding-naive policy is capped near 0.50,
# since both arms share an interventional mean.

Offline, confounded — one agent routes it. Given confounded logs, CausalMBRLAgent picks the right causal planner (back-door-adjust an observed confounder, discover the structure first, transport across a covariate shift, handle a continuous confounder, or plan a sequential regime) behind one fit → act surface, and tells you the identification it relied on.

from causalrl import CausalMBRLAgent
from causalrl.envs.suite.simpson_bandit import SimpsonBandit

env = SimpsonBandit(seed=3)                       # observed confounder Z on a back-door A <- Z -> Y
agent = CausalMBRLAgent(env.n_actions, graph=env.graph)
agent.fit(env.sample(50_000, seed=3))             # columnar {Z, A, Y} logs
agent.act({"state": 0})                           # -> 1, the interventional optimum
agent.explain()  # "CausalMBRLAgent(strategy=backdoor, adjustment_set={Z})"
# A confounding-naive marginal is fooled by Simpson's paradox and ships the worse arm.

Same class, other regimes: graph=None with variables/tiers discovers the structure first; transport=("W",) carries the policy across a covariate shift; continuous_confounder=True fits a function approximator over a continuous confounder; horizon=… plans a confounded sequential regime (the medicine DTR). The wins are confined to the confounded / offline / transfer regime by design — see the causal-MBRL results note.

What it does

Task (taxonomy) Capability Key entry points
Decision under confounding Counterfactual Thompson sampling on the MABUC CausalThompsonSampling
Confounded offline agent One front-door → back-door / discovery / transport / function-approx / sequential CausalMBRLAgent
Learned world model Fit an SCM from logs, then act in it as a Gymnasium env fit_scm, orient, counterfactual_interval
Learn the model while acting Refit the I-MEC from the agent's own experiments; Thompson-sample over structure OnlineCausalMBRL
1 — Offline→online Learn from confounded logs via causal bounds UCDTR, DOVI, DeepDeconfoundedQ
2 — Where to intervene POMIS / MIS, incl. non-manipulable variables pomis, minimal_intervention_sets
3 — Counterfactual policy Act on E[Y_do(a) | intent] CounterfactualOptimalPolicy
4 — Transportability Recover effects across domains transport_formula, transported_effect
5 — Causal discovery PC / FCI structure learning discover, CPDAG
6 — Causal imitation Imitability + confounded cloning is_imitable, CausalImitator
7 — Causal curriculum Prerequisite-ordered skill learning causal_curriculum
8 — Reward shaping Policy-invariant causal potentials causal_potential, q_learning
9 — Causal games Influence diagrams + equilibria pure_nash_equilibria, CausalGame
Identification Complete ID / gID / sID / mz; partial-ID, sensitivity & decision certificates identify_effect, manski_bounds, certify_decision

A runnable example for every row is in the Tour by Task; end-to-end notebooks are in examples/ and the Tutorials. Six task guides — five that certify a policy and one that trains an agent online — are scripts in examples/guides/ executed end to end in CI.

What the numbers say — a deliberate negative

Scored the way an RL practitioner scores things — .act() per unit, then regret against ground truth — the causal point estimates do not win on real data, and the shipped examples print that themselves rather than hiding it:

  • Twins (examples/causal_mbrl_twins.py, 11,984 pairs with both potential outcomes): our policy reaches 0.8316 survival against the per-pair oracle's 0.8747 — regret 0.0431, the worst of the learned policies, and behind the trivial constant "always the heavier twin" (0.8358).
  • LaLonde (examples/causal_mbrl_lalonde.py, priced by the NSW randomized experiment): the contextual policy enrols 73.5% of the population for $476/person of regret, where the marginal rule it is built from ("enrol everyone") leaves $0 — and its off-policy value from the observational logs, −$453/person, has the wrong sign outright.

Both examples then abstain: certify_policy refuses Twins at Γ≈1.07 and LaLonde at Γ≈1.10, and on LaLonde the randomized experiment vindicates the refusal. That is the recorded finding, and it is the positioning — the defensible edge is the decision and certificate layer, not the number. Full write-up: real-data results.

How it compares

causalrl is causal-RL-first, where the established causal libraries are estimation-first:

  • DoWhy / EconML / CausalML target treatment-effect estimation and the identify→estimate→refute workflow on i.i.d. data. They are mature, production-grade tools. causalrl instead targets sequential decision-making: intervention-set selection (POMIS), confounded offline-to-online RL, counterfactual policies, and causal curricula / shaping / games. Those are the parts of the Bareinboim taxonomy these libraries do not cover.
  • For pure graph identification it overlaps with Ananke / pgmpy / Y0.

On the RL side it is a layer, not a competitor:

  • d3rlpy (and offline-RL libraries generally) train the policy; causalrl does not reimplement any of that, and pairs with them instead. src/causalrl/scale/d3rlpy.py is the bridge in both directions — to_mdp_dataset hands a ConfoundedTrajectoryDataset to a d3rlpy algorithm, policy_actions reads the trained policy's greedy actions back, and certify_fqe wraps a fitted-Q evaluation as a certificate. pip install causalrl[scale]; see the Scale guide.
  • What it adds on top of a trained policy is the assumption those libraries' evaluators take for granted. Off-policy evaluation — importance sampling, doubly-robust, FQE — is valid only when the logged actions are unconfounded given the recorded state, and logs written by a human, a clinician or a legacy heuristic often are not. certify_policy bounds the value improvement over the logging policy under Tan's marginal sensitivity model, reports the tipping Γ at which the ship/keep decision flips, and with alpha=… gates it on a finite-sample conformal lower bound (conformal_action_value). When the bound will not carry the decision, its recommendation is abstain rather than a green light.

Use causalrl when your problem is a causal decision over time; use DoWhy/EconML when it is a treatment-effect estimate; use d3rlpy when you need the policy trained, and causalrl to decide whether to ship it.

Stability

The public API — the names exported from the top-level causalrl package — is stable and follows semantic versioning: from v1.0.0 on, breaking changes to exported names move the major version. The 0.99.x line deliberately let the surface settle in real use first; 1.0 commits to it. See Guarantees & Scope for what each method does and does not promise.

Reproducible benchmarks

uv run --extra dev python benchmarks/scbandit_report.py confounded-chain \
  --seeds 0,1,2,3,4 --steps 8000 --tail-window 2000 --n-mc 2000

The JSON report includes each seed's result plus summary uncertainty. These maintained demonstrations validate package behaviour on the stated environments; they are not general performance guarantees.

Development

uv run pytest                               # tests
uv run ruff check .                         # lint
uv run pyright src                          # types
uv run --extra docs mkdocs build --strict   # documentation

Contributions are welcome — see CONTRIBUTING.md.

Citing

If you use causalrl in research, cite the metadata in CITATION.cff and the primary source for the method you used (each is attributed inline in the Tour by Task and its source module). See Citing causalrl.

Acknowledgements

This library would not exist without the body of work it stands on. Particular thanks to:

  • Elias Bareinboim, whose 9-task taxonomy of causal reinforcement learning is the organising spine of causalrl, and whose results with collaborators are the core of nearly every slice — do-calculus completeness (with Shpitser & Pearl), transportability and selection diagrams (with Pearl), counterfactual data fusion (with Forney & Pearl), POMIS / structural causal bandits (with Lee), and causal imitation learning (with Zhang & Kumor).
  • Judea Pearl, for the do-calculus and Pearl Causal Hierarchy that make every L1 / L2 / L3 query in this library well-defined.
  • Sanghack Lee, for the reference POMIS implementation the intervention-set engine is adapted from (MIT-licensed; attribution in src/causalrl/identification/intervention_sets.py).

Other foundational references — Spirtes, Glymour & Scheines; Zhang; Manski; Tan; Koller & Milch; Ng, Harada & Russell; Bengio et al. — are cited inline at the slice that uses each.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

causalrl-3.0.0.tar.gz (1.7 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

causalrl-3.0.0-py3-none-any.whl (385.3 kB view details)

Uploaded Python 3

File details

Details for the file causalrl-3.0.0.tar.gz.

File metadata

  • Download URL: causalrl-3.0.0.tar.gz
  • Upload date:
  • Size: 1.7 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for causalrl-3.0.0.tar.gz
Algorithm Hash digest
SHA256 afd6bbbfe75d4b1b52a4f41e0a36cbaf0885938e8f1d97f458485dcf89d2d2d5
MD5 f80e3625df3e828f954cf47fbb744b31
BLAKE2b-256 b8e1b92873b047382b1d10129e43bc980c83005933218b4d44843793da6f67f6

See more details on using hashes here.

Provenance

The following attestation bundles were made for causalrl-3.0.0.tar.gz:

Publisher: publish.yml on raphaelrrcoelho/causalrl

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file causalrl-3.0.0-py3-none-any.whl.

File metadata

  • Download URL: causalrl-3.0.0-py3-none-any.whl
  • Upload date:
  • Size: 385.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for causalrl-3.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 55bb2a7ecf9cdb0e20ca67b1c617889ba29b84ac63a669f5c1a2a64bbedb7d4b
MD5 5aeb5492e6f6af2eb42fc94bf8b0bb1d
BLAKE2b-256 b1b409742dbdf39a0264ab695f90d5e6369e2867d4fe7d03cc37bc80a332cee7

See more details on using hashes here.

Provenance

The following attestation bundles were made for causalrl-3.0.0-py3-none-any.whl:

Publisher: publish.yml on raphaelrrcoelho/causalrl

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

3.0.0 This release

2 files

2.1.0

2 files

2.0.0

2 files

1.7.0

2 files

1.6.0

2 files

1.5.0

2 files

1.4.0

2 files

1.3.0

2 files

1.2.0

2 files

1.0.0

2 files

0.99.7

2 files

0.99.6

2 files

0.99.5

2 files

0.99.4

2 files

0.99.3

2 files

0.99.2

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page