CausalDemand: Benchmarking Causal Recovery of Demand from Retail Transactions and Product Text
A benchmark that scores methods on how retail sales respond to a change in price, not only on how well they forecast sales.
How the CausalDemand datasets are generated
1️⃣ Four product categories, simulated and real. Two categories are simulated: facial tissue (40 products in 731 stores) and yogurt (100 products in 721 stores). The other two are real sales of ready-to-eat cereals and snack crackers from Dominick's Finer Foods (60 products each, in 81 and 82 stores). Every category covers 156 weeks.
2️⃣ Controlled confounding. Every confounded dataset has an unconfounded twin built from the same random draws. Running a model on both shows how vulnerable it is to confounding, and any gap in its price effects can be attributed to confounding alone.
3️⃣ Valid instruments. Two instruments come with the data, so methods can also be tested on correcting the bias.
4️⃣ Equation-based and agent-based simulation. Sales are simulated in two ways: by an equation for each product and store, or by individual shoppers making their own choices. Testing on both shows whether a method works regardless of how demand is modeled.
Leaderboard
|
|
||||||||||||||||
|
|
For the simulated categories, the tables show only the confounded datasets; each value is the average over five seeds and both demand models. The real categories have no counterfactual answer key, so models are ranked by sign accuracy: the share of 10% price increases for which the model correctly predicts lower sales.
Install
CausalDemand is available on PyPI, so you can install it with pip:
pip install causaldemand
export HF_TOKEN=<your access token>
- It needs Python 3.9 or later;
rerunneeds Python 3.12 or later. - Accept its access conditions with a free Hugging Face account at
https://huggingface.co/datasets/jean-jsj/CausalDemand, then set
HF_TOKEN.
Use
causaldemand <command> <category> [<predictions_dir> | <model>] [--tissue-dose] [--tissue-switching]
Commands
download <category>: downloads every dataset of the category intocausaldemand_data/.score <category> <predictions_dir>: scores your method's predictions on every dataset.rescore <category>: re-scores the released predictions of the reference models and saves the results.rerun <category> <model>: refits reference models with the paper's settings.
Categories
tissue: facial tissue, simulated (20 datasets)yogurt: yogurt, simulated (20 datasets)cereal: ready-to-eat cereals, real sales from Dominick's (1 dataset)snack-crackers: snack crackers, real sales from Dominick's (1 dataset)all: every dataset (82), fordownloadandrescoreonly
Score your method
-
Download a category:
causaldemand download tissue -
Fit your method on each dataset and write two files from the same fitted model into a folder named like the dataset folder:
my_method/ ├── tissue_log-linear_on_seed1/ │ ├── forecast.csv │ └── scenarios.csv ├── tissue_log-linear_on_seed10/ └── ... -
Score it:
causaldemand score tissue my_method
Help
causaldemand helpshows an overview of the datasets and commands.causaldemand score --helpdescribes the input files and what to write.- The notebook
examples/baseline.ipynbwalks through a complete example.
License
- Code: Apache-2.0.
- Data on Hugging Face: CC BY-NC 4.0.
- Dominick's data: academic research only, under the terms of the James M. Kilts Center for Marketing
(
LICENSES/LicenseRef-Dominicks-Kilts.txt).- This covers the Dominick's data, the cereal and snack-cracker datasets built from them, and the brand maps in
causaldemand/dominicks_brands/(Dominick's product codes and brand names), which are not under Apache-2.0. - Work that uses them carries this acknowledgement: Dominick's data courtesy of the James M. Kilts Center for Marketing, University of Chicago Booth School of Business.
- This covers the Dominick's data, the cereal and snack-cracker datasets built from them, and the brand maps in
Citation
@unpublished{hong2026causaldemand,
title = {CausalDemand: Benchmarking Causal Recovery of Demand from Retail Transactions and Product Text},
author = {Hong, Juwon and Hwang, Minha and Shankar, Venkatesh},
note = {Working paper},
year = {2026}
}
Metadata
Release files for causaldemand 2.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| causaldemand-2.0.0.tar.gz | 345.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| causaldemand-2.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 495.9 kB
Release files / causaldemand-2.0.0.tar.gz
| Download URL | causaldemand-2.0.0.tar.gz |
|---|---|
| Size | 345.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f229f5f14915ff17cca8be754d881cb566a2bccebc94a4c530202f859997c7eb
|
|
BLAKE2b-256 checksum How to use checksums |
ec9aa67972a667a8325f5bc2b626292caa7653c376b9974e8d7ff729a25cdc6a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|
Release files / causaldemand-2.0.0-py3-none-any.whl
| Download URL | causaldemand-2.0.0-py3-none-any.whl |
|---|---|
| Size | 150.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
17dc888c027f9859794b5055fa7408cbb325ba346ee3c5f0ecf0a23782033ec2
|
|
BLAKE2b-256 checksum How to use checksums |
86fe8adbcf9caeed308fe311d0a31ca8bfd062579b058c25dc576932464df84f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|