Orphanet Explorer
A Python package for processing and merging Orphanet XML data files. This package provides tools to extract, transform, and combine data from various Orphanet XML files into a unified dataset.
Features
- Process multiple Orphanet XML file types
- Extract phenotype, functional consequences, natural history, and epidemiological data
- Merge datasets with intelligent handling of common columns
- Type-safe operations with comprehensive error handling
- Configurable output formats and locations
Installation
pip install orphanet_explorer
You can download data from https://www.orphadata.com/orphanet-scientific-knowledge-files/
Quick Start
from orphanet_explorer import OrphanetDataManager
# Initialize processor
processor = OrphanetDataManager(output_dir="output")
# Define input files
xml_files = {
"phenotype": "data/en_phenotype.xml",
"consequences": "data/en_funct_consequences.xml",
"natural_history": "data/en_nat_hist_ages.xml",
"references": "data/references.xml",
"epidemiology": "data/en_epidimiology_prev.xml"
}
# Process files and save merged dataset
merged_data = processor.process_files(
xml_files,
output_file="merged_orphanet_data.csv"
)
Use Cases
-
Medical Research
- Analyze disease phenotypes and their frequencies
- Study disease inheritance patterns
- Investigate prevalence across different populations
-
Clinical Applications
- Build reference databases for rare diseases
- Support diagnostic systems
- Track epidemiological patterns
-
Data Integration
- Combine Orphanet data with other medical databases
- Create comprehensive disease profiles
- Support machine learning models for disease classification
Basic Usage
# Process a single file type
processor = OrphanetDataManager()
root = processor.parse_xml("phenotype.xml")
phenotype_data = processor.extract_phenotype_data(root)
# Merge multiple datasets
merged_data = processor.merge_datasets([df1, df2, df3])
Contributing
We welcome contributions!
License
This project is licensed under the MIT License - see the LICENSE file for details.
Acknowledgments
- Orphanet for providing the source data
- The rare disease research community
Citation
If you use this package in your research, please cite:
@software{orphanet_processor,
author = A. Tinakoua,
title = {Orphanet Explorer: A Python Package for Processing Orphanet Data},
year = {2025},
url = {https://github.com/atinak/orphanet_explorer}
}
Metadata
Release files for orphanet-explorer 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| orphanet_explorer-0.2.0.tar.gz | 795.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| orphanet_explorer-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 810.1 kB
Release files / orphanet_explorer-0.2.0.tar.gz
| Download URL | orphanet_explorer-0.2.0.tar.gz |
|---|---|
| Size | 795.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a158c6c1109a90881bb335bc8f44ef4b03d8cc808999c2a86d55e59ce2096950
|
|
BLAKE2b-256 checksum How to use checksums |
fac4da6d417f78cad0516e43edd0d442ca0f1112998b0bf7aa7f5367ff21eca2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.9.21
|
Release files / orphanet_explorer-0.2.0-py3-none-any.whl
| Download URL | orphanet_explorer-0.2.0-py3-none-any.whl |
|---|---|
| Size | 14.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
98e3bdb1529933fa57de505797671e586c1a19a2f34349f83ada799b93d77d1e
|
|
BLAKE2b-256 checksum How to use checksums |
ebf6f32ad8f93f84b19f878a9dc81d35ba5146848b63539c2077ca54674e2a1c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.9.21
|