PyTextAD: Text Anomaly Detection in Python
PyTextAD is a Python library for detecting anomalies in text, both whole anomalous documents and the individual tokens that make them anomalous.
- 23 detectors with one interface:
fit,decision_function,predict. - Token-level detection: any detector scores every token of a document, and token scores are aggregated into document scores.
- Faithful implementations: each detector is the authors' code, or is checked numerically against it.
- Benchmark datasets with a label for every word, and evaluation at the token and document levels.
Documentation: https://pytextad.readthedocs.io
Citing PyTextAD
If you use PyTextAD, please cite it, together with the papers of the detectors you use (listed under References):
@software{cao2026pytextad,
author = {Cao, Yang},
title = {{PyTextAD}: Text Anomaly Detection in Python},
year = {2026},
url = {https://github.com/charles-cao/pytextad}
}
If you use the built-in datasets, please also cite Towards Token-Level Text Anomaly Detection (WWW 2026), listed under References.
Installation
pip install pytextad
Requires Python 3.9 or later, PyTorch 1.13 or later and transformers 4.30 or later. Install the PyTorch build that matches your CUDA version first (https://pytorch.org).
Quick start
from pytextad import SIK, TokenDetector, TokenEmbedder, align_labels
from pytextad.datasets import load_dataset
from pytextad.metrics import evaluate, format_results
ds = load_dataset("restaurant_review")
normal, anomalous = ds.normal_indices, ds.anomalous_indices
train = ds.subset(normal[:500]) # 500 normal reviews
test = ds.subset(list(normal[500:]) + list(anomalous)) # the other reviews
emb = TokenEmbedder("bert-base-uncased") # one vector per sub-word
X_train = emb.transform(train.tokens) # (vectors, word index of each sub-word)
X_test = emb.transform(test.tokens)
y_test = align_labels(test.token_labels, X_test[1]) # drops words cut off at 512 sub-words
# every sub-word is scored; a word's score is the max over its sub-words
result = evaluate(TokenDetector(SIK()), X_train, X_test, token_labels=y_test)
print(format_results({"SIK": result})) # token and document AUROC, AP, FPR95
Implemented algorithms
Text detectors take text (or token embeddings) and are trained end to end.
| Abbr | Algorithm | Year | Token scores | Ref |
|---|---|---|---|---|
| CVDD | Context Vector Data Description | 2019 | yes | [1] |
| DATE | Detecting Anomalies in Text via Self-Supervision of Transformers | 2021 | yes | [2] |
| FATE | Few-shot Anomaly Detection in Text with Deviation Learning | 2023 | no | [3] |
Embedding detectors take one vector per document or per token. Wrapped in
TokenDetector, each of them scores every sub-word, and sub-word scores are combined into
word scores.
| Abbr | Algorithm | Year | Ref |
|---|---|---|---|
| NormalizingFlow | Planar normalizing flow | 2015 | [4] |
| DAGMM | Deep Autoencoding Gaussian Mixture Model | 2018 | [5] |
| GANomaly | Adversarially trained encoder-decoder-encoder | 2018 | [6] |
| RSRAE | Robust Subspace Recovery AutoEncoder | 2020 | [7] |
| GOAD | Classification-based anomaly detection with random transformations | 2020 | [8] |
| DROCC | Distributionally Robust One-Class Classifier | 2020 | [9] |
| ICL | Internal Contrastive Learning | 2022 | [10] |
| SLAD | Scale Learning-based Anomaly Detection | 2023 | [11] |
| DTE | Diffusion Time Estimation (categorical, inverse-gamma, Gaussian, non-parametric) | 2024 | [12] |
| DDPM | Denoising diffusion model, reconstruction error | 2024 | [12] |
| MCM | Masked Cell Modeling | 2024 | [13] |
| DRL | Decomposed Representation Learning | 2025 | [14] |
| DDAE | Diffusion-Scheduled Denoising Autoencoder | 2025 | [15] |
| SIK | Simplified Isolation Kernel | 2025 | [16] |
| ADERH | Ensemble of Random Pairs of Hyperspheres | 2025 | [17] |
| TCCM | Time-Conditioned Contraction Matching | 2025 | [18] |
| TokenCore | Nearest-neighbour memory bank of token embeddings | 2026 | [19] |
Wrappers turn any embedding detector with fit and decision_function, including those of
PyOD [20], into a text detector: DocumentDetector (one embedding per document) and
TokenDetector (one embedding per token).
Datasets
Six datasets with a 0/1 label for every word, downloaded once from the Hugging Face Hub.
from pytextad.datasets import load_dataset
ds = load_dataset("restaurant_review")
ds.tokens, ds.token_labels, ds.labels # words, word labels, document labels
| Dataset | Documents | Anomalous | Anomaly |
|---|---|---|---|
sms_spam |
4,518 | 393 | injected gibberish |
restaurant_review |
1,100 | 50 | negative sentiment |
grammar_correction |
300 | 30 | grammatical errors |
hate_speech |
4,302 | 140 | hateful or offensive words |
olid |
650 | 30 | offensive words |
restaurant_review2 |
520 | 25 | negative sentiment |
License
BSD 2-Clause, except for some third-party code under its own licence (CC BY-SA 4.0, and a research-only licence for GOAD); see THIRD_PARTY_NOTICES.md.
References
[1] L. Ruff, Y. Zemlyanskiy, R. Vandermeulen, T. Schnake, M. Kloft. Self-Attentive, Multi-Context One-Class Classification for Unsupervised Anomaly Detection on Text. ACL, 2019.
[2] A. Manolache, F. Brad, E. Burceanu. DATE: Detecting Anomalies in Text via Self-Supervision of Transformers. NAACL, 2021.
[3] A. S. Das, A. Ajay, S. Saha, M. Bhuyan. Few-shot Anomaly Detection in Text with Deviation Learning. ICONIP, 2023.
[4] D. J. Rezende, S. Mohamed. Variational Inference with Normalizing Flows. ICML, 2015.
[5] B. Zong, Q. Song, M. R. Min, W. Cheng, C. Lumezanu, D. Cho, H. Chen. Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection. ICLR, 2018.
[6] S. Akcay, A. Atapour-Abarghouei, T. P. Breckon. GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training. ACCV, 2018.
[7] C.-H. Lai, D. Zou, G. Lerman. Robust Subspace Recovery Layer for Unsupervised Anomaly Detection. ICLR, 2020.
[8] L. Bergman, Y. Hoshen. Classification-Based Anomaly Detection for General Data. ICLR, 2020.
[9] S. Goyal, A. Raghunathan, M. Jain, H. V. Simhadri, P. Jain. DROCC: Deep Robust One-Class Classification. ICML, 2020.
[10] T. Shenkar, L. Wolf. Anomaly Detection for Tabular Data with Internal Contrastive Learning. ICLR, 2022.
[11] H. Xu, Y. Wang, J. Wei, S. Jian, Y. Li, N. Liu. Fascinating Supervisory Signals and Where to Find Them: Deep Anomaly Detection with Scale Learning. ICML, 2023.
[12] V. Livernoche, V. Jain, Y. Hezaveh, S. Ravanbakhsh. On Diffusion Modeling for Anomaly Detection. ICLR, 2024. arXiv:2305.18593
[13] J. Yin, Y. Qiao, Z. Zhou, X. Wang, J. Yang. MCM: Masked Cell Modeling for Anomaly Detection in Tabular Data. ICLR, 2024.
[14] H. Ye, H. Zhao, W. Fan, M. Zhou, D. Guo, Y. Chang. DRL: Decomposed Representation Learning for Tabular Anomaly Detection. ICLR, 2025.
[15] T. Sattarov, M. Schreyer, D. Borth. Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data. KDD, 2025. arXiv:2508.00758
[16] Y. Cao, S. Yang, Y. Yang, L. Qi, M. Liu. Text Anomaly Detection with Simplified Isolation Kernel. Findings of EMNLP, 2025. doi:10.18653/v1/2025.findings-emnlp.680
[17] W. Durani, C. Leiber, K. Durani, C. Plant, C. Böhm. Anomaly Detection by an Ensemble of Random Pairs of Hyperspheres. NeurIPS, 2025.
[18] Z. Li, Q. Huang, Y. Zhu, L. Yang, M. M. Amiri, N. van Stein, M. van Leeuwen. Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching. NeurIPS, 2025. arXiv:2510.18328
[19] Y. Cao, B. Yu, S. Yang, M. Liu, Y. Yang. Towards Token-Level Text Anomaly Detection. The ACM Web Conference (WWW), 2026. doi:10.1145/3774904.3792952
[20] Y. Zhao, Z. Nasrullah, Z. Li. PyOD: A Python Toolbox for Scalable Outlier Detection. JMLR, 20(96):1-7, 2019.
Metadata
Release files for pytextad 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pytextad-0.3.0.tar.gz | 104.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pytextad-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 225.9 kB
Release files / pytextad-0.3.0.tar.gz
| Download URL | pytextad-0.3.0.tar.gz |
|---|---|
| Size | 104.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
70811f602364cf029ce20ef7cc04c3164a31c09b2719e653fb704b634cbe71f5
|
|
BLAKE2b-256 checksum How to use checksums |
d530c0561c7583203e9627c7c622803c11ca1a92ac1307d5da61752c0f232f17
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.
Transparency logRelease files / pytextad-0.3.0-py3-none-any.whl
| Download URL | pytextad-0.3.0-py3-none-any.whl |
|---|---|
| Size | 121.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bdb245cadff410e3a8ee4268a4e4609f4421eddfd85d0315258b07e5eaa81d0c
|
|
BLAKE2b-256 checksum How to use checksums |
cc4a51dc9047b9897638722c7fe887b020f9ff9aac2ee0750cfe0f431bd0deb8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.
Transparency log