Skip to main content

Unified multi-touch attribution library (Markov chain + Shapley value).

Project description

do-attribution

Multi-touch attribution for Python — closed-form Markov and exact Shapley, with static and epoch-based workflows behind one package.

License: MIT Python: 3.10+

You have customer journeys that look like "Email > Social > Purchase" and a 0/1 conversion flag per journey. do-attribution turns those journeys into per-channel shares of the observed conversions using two statistical allocation models:

  • an absorbing-Markov-chain removal-effect model (MarkovAttribution, orders 1+, three removal strategies), and
  • a cooperative-game Shapley-value model (ShapleyAttribution, exact factorial weights over the observed coalition lattice).

Around the two models the package ships side-by-side model comparison (CompareAttribution), explicit conversion-endpoint handling (conversion_labels), plain-text reporting (summary()), flat DataFrame export (to_dataframes()), matplotlib/networkx visualizations, and temporal epoch analysis (TemporalMarkovAttribution, TemporalShapleyAttribution, recommend_epochs) that refits the same static models independently inside chronological epochs.

Scope guard: these are probabilistic decompositions of observed co-occurrence between channel exposure and conversion — not causal inference. A high share means the channel is statistically associated with converting paths under the model's assumptions. It does not identify the channel that caused a conversion, and it does not replace incrementality tests or randomized experiments.


Key capabilities

  • Deterministic, closed-form Markov attribution. Removal effects come analytically from the absorbing-chain fundamental matrix (I − Q)⁻¹ — no Monte Carlo, no seed, no simulation count. Identical inputs always produce identical outputs.
  • Three removal-effect definitions. removal_strategy= "fail", "detour" (signed effects), or "redirect_to_null" — three operational meanings of "remove this channel", selected explicitly instead of implied.
  • Exact Shapley values. Per-channel values use the exact factorial weight w(s, n) = s!·(n−s−1)!/n!; raw signed values, the coalition value function, and normalized shares are all returned.
  • Conservation and contract guarantees. Attribution shares sum to 1, attributed conversions conserve the observed total, 'fail' removal effects are non-negative, and absorption probabilities stay on [0, 1] — enforced by the implementation.
  • Model reporting. Every fitted model renders a compact ASCII summary() or a natural-language summary(output_format="report") — plain Python strings with bounded, distilled detail tables (signed removal effects, top transitions and drop-offs, top coalitions, path diagnostics).
  • Flat DataFrame export. Temporal models expose to_dataframes() -> dict[str, pd.DataFrame]: flat, string-keyed, Parquet/SQL-ready tables for attribution, raw model values, and sample diagnostics — no text parsing, no MultiIndex, no lazy refits.
  • Visualization. plot_attribution, plot_removal_effects, plot_transition_graph, plot_coalition_values, plot_attribution_comparison, and plot_epoch_recommendation each return a matplotlib.figure.Figure.
  • Temporal epoch analysis. Calendar buckets (freq=), explicit labels (fit_by_epoch), or data-driven boundaries (recommend_epochs + fit_recommended) — every epoch is an independent static fit; complete paths are never split across epochs.
  • Reproducible outputs. Deterministic model fits, deterministic renderers, and deterministic tie ordering — identical inputs always produce identical reports, tables, and figures.

Core runtime dependencies are numpy and pandas.


Installation

pip install do-attribution
Extra Installs Enables
viz matplotlib, networkx all plotting helpers
temporal statsmodels, ruptures global-signal epoch recommendation
excel openpyxl Excel-friendly workflows
dev all of the above + pytest running the test suite
pip install "do-attribution[viz]"
pip install "do-attribution[temporal,viz]"

Requires Python 3.10 or newer.


Quick start

A complete, runnable analysis — ten journeys, a fitted Markov model, shares, and a readable report:

from do_attribution import MarkovAttribution

journeys = [
    "Email > Social > Purchase",
    "Search > Email > Purchase",
    "Social > Purchase",
    "Email > Search",                       # no conversion
    "Search > Social > Email > Purchase",
    "Direct > Purchase",
    "Email > Social",                       # no conversion
    "Search > Email > Social > Purchase",
    "Social > Email > Purchase",
    "Search",                               # no conversion
]
conversions = [1, 1, 1, 0, 1, 1, 0, 1, 1, 0]

model = MarkovAttribution(
    order=1,
    removal_strategy="redirect_to_null",
    conversion_labels=["Purchase"],   # endpoint token, not a channel
    return_transition_matrix=True,
)
result = model.fit(journeys, conversions)

print({ch: round(share, 4) for ch, share in result["attribution"].items()})
print(f"P(conversion | start) = {result['total_conversion_probability']:.4f}")
print(model.summary())

Output (abridged — the summary continues with top transitions, drop-off bottlenecks, and path diagnostics):

{'Direct': 0.085, 'Email': 0.3912, 'Search': 0.1913, 'Social': 0.3324}
P(conversion | start) = 0.7000
Attribution Summary
Model: MarkovAttribution
Epoch source: static
Overall: 10 paths, 7 conversions, 70.00% conversion rate

Data
+-----------+-------+-------------+---------+--------------------------------------------+
| Cohort    | Paths | Conversions | Conv. % | Top channels                               |
+-----------+-------+-------------+---------+--------------------------------------------+
| All paths | 10    | 7           | 70.00%  | Email 39.12%, Social 33.24%, Search 19.13% |
+-----------+-------+-------------+---------+--------------------------------------------+

Removal effects (raw, signed) vs normalized shares
+---------+----------------+--------+
| Channel | Removal effect | Share  |
+---------+----------------+--------+
| Email   | +0.6571        | 39.12% |
| Social  | +0.5584        | 33.24% |
| Search  | +0.3214        | 19.13% |
| Direct  | +0.1429        | 8.50%  |
+---------+----------------+--------+

The Shapley model is a drop-in alternative on the same inputs:

from do_attribution import ShapleyAttribution

shapley = ShapleyAttribution(conversion_labels=["Purchase"]).fit(journeys, conversions)
print({ch: round(share, 4) for ch, share in shapley["attribution"].items()})
# {'Direct': 0.0, 'Email': 0.375, 'Search': 0.0, 'Social': 0.625}

Lists, NumPy arrays, and pandas Series all work as path and conversion inputs.


Visual examples

All figures below are outputs from the GA4 public ecommerce journey example (examples/ga4_attribution_journeys.csv — 272,792 journeys, 8 channels, Nov 2020 – Jan 2021), generated by the example notebooks with the shipped plotting helpers.

GA4 Markov attribution shares bar chart GA4 Markov transition graph
plot_attribution — Markov shares across the GA4 channels plot_transition_graph — the fitted GA4 journey chain (edges ≥ 3%)
GA4 Markov vs Shapley comparison chart GA4 epoch recommendation diagnostic plot
plot_attribution_comparison — Markov vs Shapley on the same GA4 journeys plot_epoch_recommendation — recommended epoch boundaries on the conversion-rate signal

plot_removal_effects (signed bars) and plot_coalition_values (top Shapley coalitions) complete the plotting surface; every helper returns a matplotlib.figure.Figure. The full image bank lives in examples/images/.


Supported attribution methods

Markov removal-effect attribution

MarkovAttribution(order=k) models journeys as an absorbing Markov chain (START, channels or sliding k-tuples of channels, CONVERSION, NULL) and credits each channel by its removal effect, computed in closed form. Three operational definitions of removal are provided:

Strategy Behavior
"fail" Strip the channel from every path, force conversions on paths containing it to zero, rebuild the chain, recompute P(CONVERSION | START). Effects ≥ 0.
"detour" Strip the channel while retaining original conversion flags, then refit. Effects are signed — a negative effect means removing the channel lowered modelled conversion probability.
"redirect_to_null" Redirect the removed channel's transition mass to NULL on the fitted matrix and recompute absorption analytically — the classic removal-effect definition.

Higher orders (order=2, order=3, …) use sliding tuples as states; short paths are left-padded with START. fit returns attribution, removal_effects, total_conversion_probability, and (opt-in) the transition_matrix.

Shapley coalition attribution

ShapleyAttribution uses the value function v(S) = conversions on paths whose unique channel set is exactly S and exact factorial weights. fit returns attribution, raw signed shapley_values, and the observed value_function:

shapley["value_function"]
# {('Email', 'Social'): 2, ('Email', 'Search'): 1, ('Social',): 1, ...}

Because no journey may touch every channel, the raw values can sum to zero by construction; normalization handles that regime deterministically (with a warning) instead of producing unstable shares.

Side-by-side comparison

from do_attribution import CompareAttribution

comparison = CompareAttribution(conversion_labels=["Purchase"])
comparison.fit(journeys, conversions)

comparison.attribution_frame()      #           markov  shapley
                                    # Direct    0.0263    0.000
                                    # Email     0.3947    0.375
                                    # Search    0.2018    0.000
                                    # Social    0.3772    0.625
comparison.difference_frame("markov")

The two methods answer different statistical questions and can disagree materially without either being defective — difference_frame makes the disagreement explicit. extra_models={...} plugs any additional BaseAttribution implementation into the same aligned tables, so first-touch, last-touch, or any custom baseline can sit beside the two built-in models.

Conversion-endpoint declaration

Terminal outcome tokens ("Purchase", "Sale", …) are declared, not guessed: conversion_labels=["Purchase"] strips them before channel extraction, state construction, coalition enumeration, reporting, and plotting. Labels colliding with the reserved START/CONVERSION/NULL tokens raise ValueError; conversion_labels=None treats every token as a channel.

Temporal epoch analysis

TemporalMarkovAttribution and TemporalShapleyAttribution are chronological orchestrators: each complete path is assigned to exactly one epoch by a single anchor timestamp, each epoch gets an independent static fit, and adjacent epochs share nothing — jumps at boundaries are valid outputs of a discontinuous piecewise model, not smoothing artifacts.

import pandas as pd
from do_attribution import TemporalShapleyAttribution

anchors = pd.to_datetime([
    "2025-01-05", "2025-01-12", "2025-02-08", "2025-02-17", "2025-04-03",
    "2025-04-21", "2025-07-09", "2025-07-19", "2025-10-02", "2025-10-18",
])

by_quarter = TemporalShapleyAttribution(
    freq="Q",                      # calendar buckets; per-year mappings supported
    min_samples_per_epoch=1,       # demo-sized floor
    conversion_labels=["Purchase"],
)
by_quarter.fit(journeys, anchors, conversions)

by_quarter.get_trend("Email")             # per-epoch share for one channel
tables = by_quarter.to_dataframes()       # flat DataFrames, ready for Parquet/SQL
list(tables)
# ['attribution', 'shapley_values', 'epoch_summary', 'channel_epoch_sample_sizes']

The Markov variant additionally exports removal_effects, conversion_probability, and transition_matrices tables. Three entry points cover epoch selection:

  • fit(paths, timestamps, conversions) — calendar frequency buckets;
  • fit_by_epoch(paths, epoch_labels, conversions) — explicit labels;
  • fit_recommended(paths, conversions, anchor_timestamps) — data-driven boundaries from recommend_epochs, which segments a global daily conversion-rate signal with a penalized piecewise-affine detector (weekly/multi-seasonal STL decomposition, robust second-difference scale with serial-correlation correction, PELT with a custom weighted affine cost), enforces minimum epoch width and path-count floors, and returns auditable left-closed, right-open boundaries. A Markov-only signal="transition_drift" variant segments aligned order-1 transition matrices. plot_epoch_recommendation renders the fitted segment geometry, kept and rejected knots, and per-epoch path counts.

Sample-size diagnostics (get_epoch_summary, get_channel_epoch_sample_sizes, order- and strategy-aware effective floors, and the suggest_freq calendar sanity check) flag thin epochs before you interpret an adjacent-epoch jump.

Reporting and export

Every fitted model — static, temporal, or a standalone recommendation — renders deterministic plain text:

print(model.summary())                          # compact ASCII tables
print(model.summary(output_format="report"))    # natural-language narrative

Summaries distill rather than dump: bounded top-k detail tables (signed removal effects, top transitions and drop-offs, top coalitions, path diagnostics) with the complete grids available programmatically via to_dataframes() and the get_* accessors. The reporting language is descriptive; it never labels a movement statistically significant or causal.


Detailed examples

All five notebooks run end to end on the same GA4 journey extract (examples/ga4_attribution_journeys.csv) — a transformed journey-level extract created from Google's public obfuscated GA4 ecommerce sample (Google Merchandise Store): one row per journey with an ordered channel path, a binary conversion flag, and an anchor timestamp.

Example Shows
examples/basic_markov.ipynb static Markov fit, shares, removal effects, summary, transition graph
examples/basic_shapley.ipynb Shapley shares, raw values, coalition results, coalition plot
examples/side_by_side.ipynb CompareAttribution aligned tables, differences, comparison plot
examples/temporal_markov.ipynb weekly Markov epochs, diagnostics, to_dataframes(), epoch recommendation
examples/temporal_shapley.ipynb weekly Shapley epochs, channel trends, diagnostics, epoch recommendation

Project status

do-attribution is in pre-1.0 development. The public API may change before version 0.1.


License and attribution

MIT © 2026 Aaron Goodin — see LICENSE.

The bundled example dataset and the figures generated from it derive from Google's public obfuscated GA4 ecommerce sample (Google Merchandise Store), transformed into a journey-level extract for demonstration. Attribution outputs describe statistical association under each model's assumptions; they are not causal effect estimates.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

do_attribution-0.0.1.post1.tar.gz (5.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

do_attribution-0.0.1.post1-py3-none-any.whl (94.5 kB view details)

Uploaded Python 3

File details

Details for the file do_attribution-0.0.1.post1.tar.gz.

File metadata

  • Download URL: do_attribution-0.0.1.post1.tar.gz
  • Upload date:
  • Size: 5.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for do_attribution-0.0.1.post1.tar.gz
Algorithm Hash digest
SHA256 47a046955070286ea1dc9f2c761162c43e01ecc219cd4e931b904080819839c1
MD5 726af1fa9be360e2ae8b0dcfcc5ad6ff
BLAKE2b-256 c6627e8d6bfc33e64e3be0619e7ea974532f412c84f5ca88d9e5b15fc03c4bc5

See more details on using hashes here.

File details

Details for the file do_attribution-0.0.1.post1-py3-none-any.whl.

File metadata

File hashes

Hashes for do_attribution-0.0.1.post1-py3-none-any.whl
Algorithm Hash digest
SHA256 ababde14407b3cfd4128cb5f7b145a68e71a97bd4e5b1df0527d84b299e9379b
MD5 6e87819cfdf33c3d04e6ca8cacabf29f
BLAKE2b-256 5088c826741ad9184bc91f30e1fdff1927cb65dc2653abad13115f7d3b741989

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page