CRATE: clinical records anonymisation and text extraction
Project description
# CRATE
**Clinical Records Anonymisation and Text Extraction (CRATE)**
## Purpose
- Anonymises relational databases.
- Operates a GATE natural language processing (NLP) pipeline.
- Web app for
- querying the anonymised database
- managing a consent-to-contact process
## Key directories and files
- `crate_anon/anonymise/`
- **`anonymise.py`** – core program
- `make_demo_database.py` – create a test database
- `launch_multiprocess_anonymiser.sh` – parallel processing
(multiprocess) launcher for anonymise.py
- `make_demo_database.py` – creates a demonstration database
- `test_anonymisation.py` – generates a comparison of records between
source and destination databases, to check anonymisation.
- **`crate_anon/crateweb/`** – Django web application, as above
- **`docs/`** – documentation
- `crate_anon/nlp_manager/` – NLP interface tool
- `buildjava.py` – script to compile the necessary Java source on your
machine, and create a script to test the pipeline using the ANNIE demo
GATE app.
- `CrateGatePipeline.java` – Java code to interface between
nlp_manager.py (via stdin/stdout) and the Java-based external GATE tools
(via code); must be compiled before use
- `launch_multiprocess_nlp.py` – parallel processing (multiprocess)
launcher for nlp_manager.py
- `nlp_manager.py` – core program to pipe parts of a database to a GATE
program and insert the output back into a database; uses
CrateGatePipeline.java to communicate with the NLP app
- `tools/`
- **`install_virtualenv.sh`** – creates a suitable virtualenv for CRATE
- ...
- `changelog.Debian` – Debian package changelog and general version history
- `LICENCE` – Apache license applicable to CRATE
- `README.rst` – this file
- `setup.py` – file to set up package for distribution, etc.
## Copyright/licensing
- CRATE: copyright © 2015-2016 Rudolf Cardinal (rudolf@pobox.com).
- Licensed under the Apache License, version 2.0: see LICENSE file.
- Third-party code/libraries included:
- aspects of CamAnonGatePipeline.java are based on demonstration GATE code,
copyright © University of Sheffield, and licensed under the GNU LGPL
(which license is therefore used for npl_manager/CrateGatePipeline.java;
q.v.).
**Clinical Records Anonymisation and Text Extraction (CRATE)**
## Purpose
- Anonymises relational databases.
- Operates a GATE natural language processing (NLP) pipeline.
- Web app for
- querying the anonymised database
- managing a consent-to-contact process
## Key directories and files
- `crate_anon/anonymise/`
- **`anonymise.py`** – core program
- `make_demo_database.py` – create a test database
- `launch_multiprocess_anonymiser.sh` – parallel processing
(multiprocess) launcher for anonymise.py
- `make_demo_database.py` – creates a demonstration database
- `test_anonymisation.py` – generates a comparison of records between
source and destination databases, to check anonymisation.
- **`crate_anon/crateweb/`** – Django web application, as above
- **`docs/`** – documentation
- `crate_anon/nlp_manager/` – NLP interface tool
- `buildjava.py` – script to compile the necessary Java source on your
machine, and create a script to test the pipeline using the ANNIE demo
GATE app.
- `CrateGatePipeline.java` – Java code to interface between
nlp_manager.py (via stdin/stdout) and the Java-based external GATE tools
(via code); must be compiled before use
- `launch_multiprocess_nlp.py` – parallel processing (multiprocess)
launcher for nlp_manager.py
- `nlp_manager.py` – core program to pipe parts of a database to a GATE
program and insert the output back into a database; uses
CrateGatePipeline.java to communicate with the NLP app
- `tools/`
- **`install_virtualenv.sh`** – creates a suitable virtualenv for CRATE
- ...
- `changelog.Debian` – Debian package changelog and general version history
- `LICENCE` – Apache license applicable to CRATE
- `README.rst` – this file
- `setup.py` – file to set up package for distribution, etc.
## Copyright/licensing
- CRATE: copyright © 2015-2016 Rudolf Cardinal (rudolf@pobox.com).
- Licensed under the Apache License, version 2.0: see LICENSE file.
- Third-party code/libraries included:
- aspects of CamAnonGatePipeline.java are based on demonstration GATE code,
copyright © University of Sheffield, and licensed under the GNU LGPL
(which license is therefore used for npl_manager/CrateGatePipeline.java;
q.v.).
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
crate-anon-0.17.0.tar.gz
(288.3 kB
view hashes)