Skip to main content

PyTextAD: Text Anomaly Detection in Python

PyPI Documentation Tests License

PyTextAD is a Python library for detecting anomalies in text, both whole anomalous documents and the individual tokens that make them anomalous.

  • 23 detectors with one interface: fit, decision_function, predict.
  • Token-level detection: any detector scores every token of a document, and token scores are aggregated into document scores.
  • Faithful implementations: each detector is the authors' code, or is checked numerically against it.
  • Benchmark datasets with a label for every word, and evaluation at the token and document levels.

Documentation: https://pytextad.readthedocs.io

Citing PyTextAD

If you use PyTextAD, please cite it, together with the papers of the detectors you use (listed under References):

@software{cao2026pytextad,
  author = {Cao, Yang},
  title  = {{PyTextAD}: Text Anomaly Detection in Python},
  year   = {2026},
  url    = {https://github.com/charles-cao/pytextad}
}

If you use the built-in datasets, please also cite Towards Token-Level Text Anomaly Detection (WWW 2026), listed under References.

Installation

pip install pytextad

Requires Python 3.9 or later, PyTorch 1.13 or later and transformers 4.30 or later. Install the PyTorch build that matches your CUDA version first (https://pytorch.org).

Quick start

from pytextad import SIK, TokenDetector, TokenEmbedder, align_labels
from pytextad.datasets import load_dataset
from pytextad.metrics import evaluate, format_results

ds = load_dataset("restaurant_review")
normal, anomalous = ds.normal_indices, ds.anomalous_indices
train = ds.subset(normal[:500])                         # 500 normal reviews
test = ds.subset(list(normal[500:]) + list(anomalous))  # the other reviews

emb = TokenEmbedder("bert-base-uncased")              # one vector per sub-word
X_train = emb.transform(train.tokens)                 # (vectors, word index of each sub-word)
X_test = emb.transform(test.tokens)
y_test = align_labels(test.token_labels, X_test[1])  # drops words cut off at 512 sub-words

# every sub-word is scored; a word's score is the max over its sub-words
result = evaluate(TokenDetector(SIK()), X_train, X_test, token_labels=y_test)
print(format_results({"SIK": result}))     # token and document AUROC, AP, FPR95

Implemented algorithms

Text detectors take text (or token embeddings) and are trained end to end.

Abbr Algorithm Year Token scores Ref
CVDD Context Vector Data Description 2019 yes [1]
DATE Detecting Anomalies in Text via Self-Supervision of Transformers 2021 yes [2]
FATE Few-shot Anomaly Detection in Text with Deviation Learning 2023 no [3]

Embedding detectors take one vector per document or per token. Wrapped in TokenDetector, each of them scores every sub-word, and sub-word scores are combined into word scores.

Abbr Algorithm Year Ref
NormalizingFlow Planar normalizing flow 2015 [4]
DAGMM Deep Autoencoding Gaussian Mixture Model 2018 [5]
GANomaly Adversarially trained encoder-decoder-encoder 2018 [6]
RSRAE Robust Subspace Recovery AutoEncoder 2020 [7]
GOAD Classification-based anomaly detection with random transformations 2020 [8]
DROCC Distributionally Robust One-Class Classifier 2020 [9]
ICL Internal Contrastive Learning 2022 [10]
SLAD Scale Learning-based Anomaly Detection 2023 [11]
DTE Diffusion Time Estimation (categorical, inverse-gamma, Gaussian, non-parametric) 2024 [12]
DDPM Denoising diffusion model, reconstruction error 2024 [12]
MCM Masked Cell Modeling 2024 [13]
DRL Decomposed Representation Learning 2025 [14]
DDAE Diffusion-Scheduled Denoising Autoencoder 2025 [15]
SIK Simplified Isolation Kernel 2025 [16]
ADERH Ensemble of Random Pairs of Hyperspheres 2025 [17]
TCCM Time-Conditioned Contraction Matching 2025 [18]
TokenCore Nearest-neighbour memory bank of token embeddings 2026 [19]

Wrappers turn any embedding detector with fit and decision_function, including those of PyOD [20], into a text detector: DocumentDetector (one embedding per document) and TokenDetector (one embedding per token).

Datasets

Six datasets with a 0/1 label for every word, downloaded once from the Hugging Face Hub.

from pytextad.datasets import load_dataset
ds = load_dataset("restaurant_review")
ds.tokens, ds.token_labels, ds.labels      # words, word labels, document labels
Dataset Documents Anomalous Anomaly
sms_spam 4,518 393 injected gibberish
restaurant_review 1,100 50 negative sentiment
grammar_correction 300 30 grammatical errors
hate_speech 4,302 140 hateful or offensive words
olid 650 30 offensive words
restaurant_review2 520 25 negative sentiment

License

BSD 2-Clause, except for some third-party code under its own licence (CC BY-SA 4.0, and a research-only licence for GOAD); see THIRD_PARTY_NOTICES.md.

References

[1] L. Ruff, Y. Zemlyanskiy, R. Vandermeulen, T. Schnake, M. Kloft. Self-Attentive, Multi-Context One-Class Classification for Unsupervised Anomaly Detection on Text. ACL, 2019.

[2] A. Manolache, F. Brad, E. Burceanu. DATE: Detecting Anomalies in Text via Self-Supervision of Transformers. NAACL, 2021.

[3] A. S. Das, A. Ajay, S. Saha, M. Bhuyan. Few-shot Anomaly Detection in Text with Deviation Learning. ICONIP, 2023.

[4] D. J. Rezende, S. Mohamed. Variational Inference with Normalizing Flows. ICML, 2015.

[5] B. Zong, Q. Song, M. R. Min, W. Cheng, C. Lumezanu, D. Cho, H. Chen. Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection. ICLR, 2018.

[6] S. Akcay, A. Atapour-Abarghouei, T. P. Breckon. GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training. ACCV, 2018.

[7] C.-H. Lai, D. Zou, G. Lerman. Robust Subspace Recovery Layer for Unsupervised Anomaly Detection. ICLR, 2020.

[8] L. Bergman, Y. Hoshen. Classification-Based Anomaly Detection for General Data. ICLR, 2020.

[9] S. Goyal, A. Raghunathan, M. Jain, H. V. Simhadri, P. Jain. DROCC: Deep Robust One-Class Classification. ICML, 2020.

[10] T. Shenkar, L. Wolf. Anomaly Detection for Tabular Data with Internal Contrastive Learning. ICLR, 2022.

[11] H. Xu, Y. Wang, J. Wei, S. Jian, Y. Li, N. Liu. Fascinating Supervisory Signals and Where to Find Them: Deep Anomaly Detection with Scale Learning. ICML, 2023.

[12] V. Livernoche, V. Jain, Y. Hezaveh, S. Ravanbakhsh. On Diffusion Modeling for Anomaly Detection. ICLR, 2024. arXiv:2305.18593

[13] J. Yin, Y. Qiao, Z. Zhou, X. Wang, J. Yang. MCM: Masked Cell Modeling for Anomaly Detection in Tabular Data. ICLR, 2024.

[14] H. Ye, H. Zhao, W. Fan, M. Zhou, D. Guo, Y. Chang. DRL: Decomposed Representation Learning for Tabular Anomaly Detection. ICLR, 2025.

[15] T. Sattarov, M. Schreyer, D. Borth. Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data. KDD, 2025. arXiv:2508.00758

[16] Y. Cao, S. Yang, Y. Yang, L. Qi, M. Liu. Text Anomaly Detection with Simplified Isolation Kernel. Findings of EMNLP, 2025. doi:10.18653/v1/2025.findings-emnlp.680

[17] W. Durani, C. Leiber, K. Durani, C. Plant, C. Böhm. Anomaly Detection by an Ensemble of Random Pairs of Hyperspheres. NeurIPS, 2025.

[18] Z. Li, Q. Huang, Y. Zhu, L. Yang, M. M. Amiri, N. van Stein, M. van Leeuwen. Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching. NeurIPS, 2025. arXiv:2510.18328

[19] Y. Cao, B. Yu, S. Yang, M. Liu, Y. Yang. Towards Token-Level Text Anomaly Detection. The ACM Web Conference (WWW), 2026. doi:10.1145/3774904.3792952

[20] Y. Zhao, Z. Nasrullah, Z. Li. PyOD: A Python Toolbox for Scalable Outlier Detection. JMLR, 20(96):1-7, 2019.

Metadata

Release files for pytextad 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pytextad 0.3.0
File Size Uploaded
pytextad-0.3.0.tar.gz 104.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pytextad 0.3.0
File Interpreter ABI Platform
pytextad-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 225.9 kB

Release files / pytextad-0.3.0.tar.gz

Download URL pytextad-0.3.0.tar.gz
Size 104.9 kB
Tags Source
SHA-256 checksum
How to use checksums
70811f602364cf029ce20ef7cc04c3164a31c09b2719e653fb704b634cbe71f5
BLAKE2b-256 checksum
How to use checksums
d530c0561c7583203e9627c7c622803c11ca1a92ac1307d5da61752c0f232f17
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / pytextad-0.3.0-py3-none-any.whl

Download URL pytextad-0.3.0-py3-none-any.whl
Size 121.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bdb245cadff410e3a8ee4268a4e4609f4421eddfd85d0315258b07e5eaa81d0c
BLAKE2b-256 checksum
How to use checksums
cc4a51dc9047b9897638722c7fe887b020f9ff9aac2ee0750cfe0f431bd0deb8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page