Skip to main content

PyCTAKES 🏥

Open Source Python-native Clinical NLP Framework

License: MIT Python 3.8+ GitHub issues GitHub stars GitHub forks Contributions welcome

🚀 A modern, open source clinical NLP framework that mirrors and extends Apache cTAKES functionality in pure Python. Drop-in replacement with superior usability, extensibility, and performance.

PyCTAKES transforms clinical text processing by providing a 100% open source, Python-native alternative to Apache cTAKES. Built by the community, for the community - no vendor lock-in, no licensing fees, just powerful clinical NLP tools that anyone can use, modify, and contribute to.


🌟 Why Choose PyCTAKES?

🔓 Fully Open Source

  • MIT License - Free for commercial & research use
  • Transparent development - All code, issues, and discussions public
  • Community-driven - Shaped by real user needs
  • No vendor lock-in - Own your clinical NLP pipeline

⚡ Modern & Fast

  • Pure Python - No Java dependencies
  • pip installable - Get started in seconds
  • Multiple backends - spaCy, Stanza, rule-based
  • Production ready - Optimized for real-world use

🏥 Clinical-First Design

  • Medical expertise built-in - Clinical abbreviations, sections, terminology
  • cTAKES compatibility - Drop-in replacement for existing workflows
  • Comprehensive NLP - Tokenization → UMLS mapping
  • Assertion detection - Negation, uncertainty, temporal context

🔧 Developer Friendly

  • Clean Python APIs - Intuitive and well-documented
  • Modular architecture - Use only what you need
  • Extensible framework - Easy to add custom annotators
  • Rich ecosystem - Integrates with pandas, spaCy, transformers

🚀 Quick Start

Installation

pip install pytakes

30-Second Demo

import pytakes

# Create pipeline
pipeline = pytakes.create_default_pipeline()

# Process clinical text
clinical_note = """
Patient is a 65-year-old male with diabetes and hypertension.
He denies chest pain but reports shortness of breath.
Current medications: metformin 500mg BID, lisinopril 10mg daily.
"""

result = pipeline.process_text(clinical_note)

# Explore results
print(f"Found {len(result.entities)} clinical entities:")
for entity in result.entities[:3]:
    assertion = entity.assertion
    print(f"  • {entity.text} ({entity.label})")
    print(f"    → {assertion.polarity}, {assertion.uncertainty}")

Output:

Found 8 clinical entities:
  • diabetes (CONDITION)
    → POSITIVE, CERTAIN
  • hypertension (CONDITION)  
    → POSITIVE, CERTAIN
  • chest pain (SYMPTOM)
    → NEGATIVE, CERTAIN

📊 Performance & Features

⚡ Blazing Fast Performance

  • Basic Pipeline: 39 annotations in 0.010s
  • Fast Pipeline: 36 annotations in 0.001s
  • Full Clinical Note: 81 annotations in 0.504s

🎯 Comprehensive Clinical NLP

Feature Description Status
Sentence Segmentation Clinical-aware sentence boundary detection ✅
Tokenization Advanced tokenization with POS tagging ✅
Section Detection Chief Complaint, History, Medications, Assessment, etc. ✅
Named Entity Recognition Medications, conditions, procedures, anatomy ✅
Assertion Detection Negation, uncertainty, temporal, experiencer ✅
UMLS Concept Mapping CUI normalization and semantic types ✅
Relation Extraction Temporal and dosage relationships 🔄 v1.1
REST API Service FastAPI deployment wrapper 🔄 v1.1

🔧 Three Pipeline Types

# Full-featured (highest accuracy)
pipeline = pytakes.create_default_pipeline()

# Speed-optimized (fastest processing)  
pipeline = pytakes.create_fast_pipeline()

# Minimal (basic entity extraction)
pipeline = pytakes.create_basic_pipeline()

💻 Command Line Interface

# Process single file
pytakes process note.txt --output results.json

# Batch processing
pytakes process notes/*.txt --output-dir results/

# Different pipelines and formats
pytakes process note.txt --pipeline fast --format xml
pytakes process note.txt --config custom_config.json

🤝 Open Source Community

👥 Lead Contributors

  • Sonish Sivarajkumar - Lead Maintainer & Creator
    • Clinical NLP researcher and software engineer
    • Apache cTAKES community member
    • Python & healthcare technology enthusiast

🌍 Join Our Community

We're building the future of clinical NLP together! Whether you're a:

  • 👩‍⚕️ Clinician - Help us understand real-world clinical text challenges
  • 👨‍💻 Developer - Contribute code, fix bugs, or add new features
  • 🔬 Researcher - Share use cases, benchmarks, and domain expertise
  • 📚 Technical Writer - Improve documentation and tutorials
  • 🎨 Designer - Enhance user experience and visualization

Everyone is welcome! Check out our Contributing Guide to get started.

📈 Community Stats

  • Contributors: Growing community of clinical NLP enthusiasts
  • Issues: Active issue tracking and feature requests
  • Discussions: Technical discussions and use case sharing
  • Releases: Regular updates with new features and improvements

🎯 Ways to Contribute

🐛 Report Issues

  • Bug reports
  • Feature requests
  • Documentation issues
  • Performance problems

💡 Share Ideas

  • New annotators
  • Pipeline improvements
  • Integration suggestions
  • Use case examples

🔧 Code Contributions

  • Bug fixes
  • New features
  • Performance optimizations
  • Test improvements

📖 Documentation

  • API documentation
  • Tutorials & guides
  • Example notebooks
  • Translation support

📚 Documentation & Resources


🗺️ Roadmap

🎯 v1.0 (Current) - Foundation

  • ✅ Core pipeline architecture
  • ✅ Clinical text processing (tokenization, NER, assertion)
  • ✅ UMLS concept mapping framework
  • ✅ CLI and Python APIs
  • ✅ Comprehensive documentation

⚡ v1.1 (Next) - Enhancement

  • 🔄 Enhanced UMLS integration (QuickUMLS)
  • 🔄 Relation extraction (temporal, dosage)
  • 🔄 REST API service wrapper
  • 🔄 Docker containers & deployment guides
  • 🔄 Performance optimizations

🚀 v2.0 (Future) - Intelligence

  • 🔮 LLM integration for disambiguation
  • 🔮 Active learning capabilities
  • 🔮 Advanced relation extraction
  • 🔮 Real-time processing pipelines
  • 🔮 Federated learning support

🏆 Why Open Source Matters in Healthcare

Healthcare technology should be:

  • 🔍 Transparent - Auditable algorithms for patient safety
  • 🤝 Collaborative - Shared knowledge accelerates progress
  • ♿ Accessible - No barriers to life-saving technology
  • 🔧 Customizable - Adaptable to diverse clinical environments
  • 📈 Sustainable - Community-driven long-term maintenance

PyTAKES embodies these principles by providing enterprise-grade clinical NLP capabilities as a truly open source project. No hidden costs, no vendor dependencies, just powerful tools for advancing healthcare through technology.


📄 License & Citation

📜 License

PyTAKES is released under the MIT License - see LICENSE for details.

Copyright (c) 2025 Sonish Sivarajkumar and Contributors

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
[Full license text in LICENSE file]

📝 Citation

If you use PyTAKES in your research, please cite:

@software{pytakes2025,
  title={PyTAKES: Open Source Python-native Clinical NLP Framework},
  author={Sivarajkumar, Sonish and Contributors},
  year={2025},
  url={https://github.com/sonishsivarajkumar/PyTAKES},
  version={1.0.0}
}

🙏 Acknowledgments

PyTAKES builds upon the excellent work of:

  • Apache cTAKES - Pioneering clinical NLP framework
  • spaCy & Stanza - Modern NLP processing libraries
  • Clinical NLP Community - Researchers and practitioners advancing the field
  • Open Source Contributors - Everyone who helps make this project better

🚀 Get Started Today!

# Install PyTAKES
pip install pytakes

# Clone the repository
git clone https://github.com/sonishsivarajkumar/PyTAKES.git
cd PyTAKES

# Try the examples
python examples/comprehensive_demo.py

Join us in revolutionizing clinical NLP! 🎉

⭐ Star this repo | 📚 Read the docs | 🤝 Contribute | 💬 Discuss

Metadata

Release files for pyctakes 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyctakes 1.0.0
File Size Uploaded
pyctakes-1.0.0.tar.gz 37.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyctakes 1.0.0
File Interpreter ABI Platform
pyctakes-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 69.8 kB

Release files / pyctakes-1.0.0.tar.gz

Download URL pyctakes-1.0.0.tar.gz
Size 37.5 kB
Tags Source
SHA-256 checksum
How to use checksums
360f2faa2fa1d9404dbfd697d73b84699872e05e48ba27e9ca5b31029b37c9cd
BLAKE2b-256 checksum
How to use checksums
9b91053264c4a8d50321c395672f25aa1c26373e5a2aa8f07a00b7d26fe2539f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.16

Release files / pyctakes-1.0.0-py3-none-any.whl

Download URL pyctakes-1.0.0-py3-none-any.whl
Size 32.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4b22bc146815495c468a5e6e2c608f29acfdfacdcbefedb284ebc586473196d5
BLAKE2b-256 checksum
How to use checksums
c63e3586e01faf3b3aa88439dd55b1abd03da78d54a86f8611c49b37e71fcb9c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.16

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page