🔒 READI - Risk Evaluation and De-Identification
Privacy-preserving AI made simple - A comprehensive toolkit for data privacy risk assessment and de-identification in Python-based ML pipelines.
READI augments the functionalities provided by IBM Data Privacy Toolkit, offering state-of-the-art capabilities for detecting Personal and Sensitive Information in unstructured documents. Built for modern compliance frameworks and AI model training workflows.
✨ Features
- 🎯 Advanced PII Detection - Identify personal and sensitive information across multiple data types
- 🔄 Seamless Integration - Low-effort integration with existing ML pipelines
- 📊 Structured & Unstructured Data - Support for both data formats
- 🌐 REST API - Easy-to-use HTTP interface for remote processing
- 🧪 Extensible Framework - Modular design for custom privacy requirements
- 📝 Comprehensive Examples - Jupyter notebooks with real-world use cases
🚀 Quick Start
Prerequisites
- Python 3.11 or higher
- Git with git-lfs support (for large files >50 MB)
- uv (recommended) - A fast Python package installer
Installation
Recommended: Using uv (10-100x faster)
# Install uv if you haven't already
curl -LsSf https://astral.sh/uv/install.sh | sh
# Create and activate virtual environment
uv venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# Install READI
uv pip install git+https://github.com/IBM/READI.git
Standard Installation with pip:
pip install git+https://github.com/IBM/READI.git
Clone Repository:
git clone https://github.com/IBM/READI.git
cd READI
# With uv (recommended)
uv pip install -e .
# Or with pip
pip install -e .
💻 Development Setup
For contributors and developers:
Recommended: Using uv
# Install in editable mode with development dependencies
uv pip install -e .
uv pip install -r requirements-dev.txt
# Set up pre-commit hooks (recommended)
pre-commit install
Alternative: Using pip
# Install in editable mode with development dependencies
pip install -e .
pip install -r requirements-dev.txt
# Set up pre-commit hooks (recommended)
pre-commit install
This installs the project in editable mode along with development tools (pytest, ruff, bandit, etc.).
💡 Tip: Using
uvprovides significantly faster dependency resolution and installation compared to traditionalpip.
🌐 REST API Usage
READI provides a simple REST API for remote processing.
Setup
# Install with REST API support
pip install -e '.[rest]'
# Start the server
uvicorn risk_assessment.entry_points.rest.api:app
Example Request
curl -H 'Content-Type: application/json' \
http://localhost:8000/detect_phi \
--data-raw '{"text":"My text with email: john@gmail.com"}'
The API will be available at http://localhost:8000 with interactive documentation at /docs.
📚 Examples & Tutorials
Explore our comprehensive Jupyter notebooks in the notebooks/ directory:
| Notebook | Description |
|---|---|
| Unstructured Data Classification | General overview of READI API for free-text processing |
| Structured Data Classification | Working with tabular and structured datasets |
| Unstructured Data Masking | Applying masking actions (redaction, tagging, hash) after PII classification |
📖 Documentation
For detailed documentation, API references, and advanced usage patterns, please visit our documentation portal (coming soon).
🤝 Contributing
We welcome contributions! Please see our Contributing Guidelines for details on:
- Code style and standards
- Testing requirements
- Pull request process
- Development workflow
📄 License
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
📌 How to Cite
If you use READI in academic work, please cite the most relevant publication from the references below. A general citation entry is:
@software{readi_ibm,
title = {READI: Risk Evaluation and De-Identification},
author = {Stefano Braghin and Liubov Nedoshivina and Anisa Halimi and Naoise Holohan and Kieran Fraser},
year = {2026},
url = {https://github.com/IBM/READI}
}
When your usage specifically relates to unstructured document de-identification, prefer citing:
@article{nedoshivina2024pragmatic,
title = {Pragmatic De-Identification of Cross-Domain Unstructured Documents: A Utility-Preserving Approach with Relation Extraction Filtering},
author = {Liubov Nedoshivina and Anisa Halimi and Joa Bettencourt-Silva and Stefano Braghin},
journal = {AMIA Summits on Translational Science Proceedings},
volume = {2024},
pages = {85},
year = {2024}
}
📚 Academic References
READI is built on years of privacy research. Key publications:
-
Nedoshivina, L., Halimi, A., Bettencourt-Silva, J., & Braghin, S. (2024). Pragmatic De-Identification of Cross-Domain Unstructured Documents: A Utility-Preserving Approach with Relation Extraction Filtering. AMIA Summits on Translational Science Proceedings, 2024, 85.
-
Pachilakis, M., Antonatos, S., Levacher, K., & Braghin, S. (2020). PrivLeAD: Privacy Leakage Detection on the Web. Intelligent Systems and Applications. IntelliSys 2020. Advances in Intelligent Systems and Computing, vol 1250. Springer, Cham. DOI: 10.1007/978-3-030-55180-3_32
-
Braghin, S., Bettencourt-Silva, J. H., Levacher, K., & Antonatos, S. (2019). An Extensible De-Identification Framework for Privacy Protection of Unstructured Health Information: Creating Sustainable Privacy Infrastructures. MEDINFO 2019: Health and Wellbeing e-Networks for All (pp. 1140-1144). IOS Press. DOI: 10.3233/SHTI190404
-
Antonatos, S., Braghin, S., Holohan, N., Gkoufas, Y., & Mac Aonghusa, P. (2018). PRIMA: An End-to-End Framework for Privacy at Scale. 2018 IEEE 34th International Conference on Data Engineering (ICDE), pp. 1531-1542. DOI: 10.1109/ICDE.2018.00171
-
Gkoulalas-Divanis, A., & Braghin, S. (2016). IPV: A system for identifying privacy vulnerabilities in datasets. IBM Journal of Research and Development, vol. 60, no. 4, pp. 14:1-14:10. DOI: 10.1147/JRD.2016.2576818
-
Gkoulalas-Divanis, A., Braghin, S., & Antonatos, S. (2016). FPVI: A scalable method for discovering privacy vulnerabilities in microdata. 2016 IEEE International Smart Cities Conference (ISC2), pp. 1-8. DOI: 10.1109/ISC2.2016.7580849
-
Gkoulalas-Divanis, A., & Braghin, S. (2015). Efficient algorithms for identifying privacy vulnerabilities. 2015 IEEE First International Smart Cities Conference (ISC2), pp. 1-8. DOI: 10.1109/ISC2.2015.7366170
🙏 Acknowledgment
This project is partly supported by the Innovative Health Initiative Joint Undertaking (IHI JU) under Grant Agreement No. 101172997 – SEARCH, and by the European Union’s Horizon research and innovation programme under Grant Agreement No. 101298664 - RegulAIze.
💬 Support & Community
- 🐛 Issues: GitHub Issues
- 💡 Discussions: GitHub Discussions
- 📧 Contact: For enterprise support, please contact the IBM Research team
Built with ❤️ by IBM Research
Metadata
Release files for readi-privacy 0.1.7
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| readi_privacy-0.1.7.tar.gz | 16.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| readi_privacy-0.1.7-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 29.7 MB
Release files / readi_privacy-0.1.7.tar.gz
| Download URL | readi_privacy-0.1.7.tar.gz |
|---|---|
| Size | 16.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8df41dac418dcded3231beffb452e1e58e2602ab4b8e8050610ea867e334e9c3
|
|
BLAKE2b-256 checksum How to use checksums |
ed96cffb74f5043527ad35fe1da1927afa88efa0f47d7da8e8af45e462453c8f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency logRelease files / readi_privacy-0.1.7-py3-none-any.whl
| Download URL | readi_privacy-0.1.7-py3-none-any.whl |
|---|---|
| Size | 13.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
87d37bde757a396d7e708048b39844e83cac4cfa95a48fe3e65c95edaee7c87d
|
|
BLAKE2b-256 checksum How to use checksums |
15cfd0f32e5422b2803bd45f53da2d08fbd5ce2a279fc929c9490fd6b04d68b6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency log