Skip to main content

Text Explainability logo

A generic explainability architecture for explaining text machine learning models

PyPI Downloads Python_version Lint, Security & Tests License Documentation Status Code style: black DOI


text_explainability provides a generic architecture from which well-known state-of-the-art explainability approaches for text can be composed. This modular architecture allows components to be swapped out and combined, to quickly develop new types of explainability approaches for (natural language) text, or to improve a plethora of approaches by improving a single module.

Several example methods are included, which provide local explanations (explaining the prediction of a single instance, e.g. LIME and SHAP) or global explanations (explaining the dataset, or model behavior on the dataset, e.g. TokenFrequency and MMDCritic). By replacing the default modules (e.g. local data generation, global data sampling or improved embedding methods), these methods can be improved upon or new methods can be introduced.

© Marcel Robeer, 2021

Quick tour

Local explanation: explain a models' prediction on a given sample, self-provided or from a dataset.

from text_explainability import LIME, LocalTree

# Define sample to explain
sample = 'Explain why this is positive and not negative!'

# LIME explanation (local feature importance)
LIME().explain(sample, model).scores

# List of local rules, extracted from tree
LocalTree().explain(sample, model).rules

Global explanation: explain the whole dataset (e.g. train set, test set), and what they look like for the ground-truth or predicted labels.

from text_explainability import import_data, TokenFrequency, MMDCritic

# Import dataset
env = import_data('./datasets/test.csv', data_cols=['fulltext'], label_cols=['label'])

# Top-k most frequent tokens per label
TokenFrequency(env.dataset).explain(labelprovider=env.labels, explain_model=False, k=3)

# 2 prototypes and 1 criticisms for the dataset
MMDCritic(env.dataset)(n_prototypes=2, n_criticisms=1)

Installation

See the installation instructions for an extended installation guide.

Method Instructions
pip Install from PyPI via pip3 install text_explainability. To speed up the explanation generation process use pip3 install text_explainability[fast].
Local Clone this repository and install via pip3 install -e . or locally run python3 setup.py install.

Documentation

Full documentation of the latest version is provided at https://text-explainability.readthedocs.io/.

Example usage

See example usage to see an example of how the package can be used, or run the lines in example_usage.py to do explore it interactively.

Explanation methods included

text_explainability includes methods for model-agnostic local explanation and global explanation. Each of these methods can be fully customized to fit the explainees' needs.

Type Explanation method Description Paper/link
Local explanation LIME Calculate feature attribution with Local Intepretable Model-Agnostic Explanations (LIME). [Ribeiro2016], interpretable-ml/lime
KernelSHAP Calculate feature attribution with Shapley Additive Explanations (SHAP). [Lundberg2017], interpretable-ml/shap
LocalTree Fit a local decision tree around a single decision. [Guidotti2018]
LocalRules Fit a local sparse set of label-specific rules using SkopeRules. github/skope-rules
FoilTree Fit a local contrastive/counterfactual decision tree around a single decision. [Robeer2018]
BayLIME Bayesian extension of LIME for include prior knowledge and more consistent explanations. [Zhao201]
Global explanation TokenFrequency Show the top-k number of tokens for each ground-truth or predicted label.
TokenInformation Show the top-k token mutual information for a dataset or model. wikipedia/mutual_information
KMedoids Embed instances and find top-n prototypes (can also be performed for each label using LabelwiseKMedoids). interpretable-ml/prototypes
MMDCritic Embed instances and find top-n prototypes and top-n criticisms (can also be performed for each label using LabelwiseMMDCritic). [Kim2016], interpretable-ml/prototypes

Releases

text_explainability is officially released through PyPI.

See CHANGELOG.md for a full overview of the changes for each version.

Extensions

Text sensitivity logo

text_explainability can be extended to also perform sensitivity testing, checking for machine learning model robustness and fairness. The text_sensitivity package is available through PyPI and fully documented at https://text-sensitivity.rtfd.io/.

Citation

@misc{text_explainability,
  title = {Python package text\_explainability},
  author = {Marcel Robeer},
  howpublished = {\url{https://github.com/MarcelRobeer/text_explainability}}
  doi = {10.5281/zenodo.14192126},
  year = {2021}
}

Maintenance

Contributors

Todo

Tasks yet to be done:

  • Implement local post-hoc explanations:
    • Implement Anchors
  • Implement global post-hoc explanations:
    • Representative subset
  • Add support for regression models
  • More complex data augmentation
    • Top-k replacement (e.g. according to LM / WordNet)
    • Tokens to exclude from being changed
    • Bag-of-words style replacements
  • Add rule-based return type
  • Write more tests

Credits

Metadata

Release files for text-explainability 0.7.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for text-explainability 0.7.2
File Size Uploaded
text_explainability-0.7.2.tar.gz 50.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for text-explainability 0.7.2
File Interpreter ABI Platform
text_explainability-0.7.2-py3-none-any.whl Python 3 none any Details

Total release size: 98.1 kB

Release files / text_explainability-0.7.2.tar.gz

Download URL text_explainability-0.7.2.tar.gz
Size 50.4 kB
Tags Source
SHA-256 checksum
How to use checksums
b2f47167f7fff988be307ac752dd28765830344d670d35db1ad3315bc5b70a31
BLAKE2b-256 checksum
How to use checksums
5f1d75531977e539e7e588cbc71150bf6ffd13bab1477bbeb98509e5b0f52ec1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.3

Release files / text_explainability-0.7.2-py3-none-any.whl

Download URL text_explainability-0.7.2-py3-none-any.whl
Size 47.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0156503d932227f8553f82efe6d7e04504cabdfd0fa217b0c5ea37751f9f3528
BLAKE2b-256 checksum
How to use checksums
01454fb32f4e409f848f2578e1ef2d178dad1ad6d7a92e37c9613b527bc68a52
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.3

Release history Release notifications | RSS feed

This release

0.7.2 This release

2 release files

0.7.0

2 release files

0.6.7

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.8

2 release files

0.5.7

3 release files

0.5.6

3 release files

0.5.5

3 release files

0.5.4

2 release files

0.5.3

3 release files

0.5.2

2 release files

0.5.1

3 release files

0.5.0

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

3 release files

0.4.2

3 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.8

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

3 release files

0.3.0

2 release files

0.2

2 release files

0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page