Medical Triage & Clinical RAG Engine
Project description
⭐ If Ella's RAG-based medical triage architecture gave you ideas — a star helps other health-AI builders find it. Takes 2 seconds.
ELLA
Medical Triage & Clinical RAG Engine
Ella is a production-grade Retrieval-Augmented Generation (RAG) system purpose-built for medical triage. She ingests 90,000+ clinical text chunks, embeds them via NVIDIA NIM, stores them in Pinecone, and retrieves context-grounded answers through a multi-stage pipeline — eliminating hallucinations in healthcare workflows.
Install • Quick Start • Architecture • Live Demo • Benchmark
Install
pip install ella-sdk
Quick Start
from ella_medical import Ella
client = Ella()
response = client.query("What are the symptoms of a heart attack?")
print(response.intent) # "TRIAGE"
print(response.response) # Grounded clinical response
Usage
Basic Query
from ella_medical import Ella
client = Ella()
response = client.query("What are the symptoms of a heart attack?")
print(response.intent) # "TRIAGE"
print(response.priority) # Priority level
print(response.thought_process) # Router's reasoning
print(response.response) # Ella's response
print(response.retrieved_context) # Retrieved medical documents
Multi-Turn Conversation
from ella_medical import Ella
client = Ella()
# First message
r1 = client.query("I have chest pain")
# Follow-up
r2 = client.query(
"What about treatment options?",
history=f"Patient: I have chest pain\nElla: {r1.response}"
)
print(r2.response)
Context Manager
from ella_medical import Ella
with Ella() as client:
response = client.query("What are the symptoms of diabetes?")
print(response.response)
Response Object
@dataclass
class QueryResponse:
intent: str # EMERGENCY | TRIAGE | BOOKING | GENERAL_INFO | CLOSING
priority: str # Priority level
thought_process: str # Router's reasoning
justification: str # Clinical justification
response: str # Ella's response
retrieved_context: str # Retrieved medical documents
Live Demo
Architecture
┌─────────────────────────────────────────────────────────────────────┐
│ USER'S INPUT │
└────────────────────────────┬────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────┐
│ INTENT ROUTER (Groq llama-3.1-8b-instant + Pydantic Schema) │
│ Classifies: EMERGENCY │ TRIAGE │ BOOKING │ GENERAL_INFO │ CLOSING │
└────────────────────────────┬────────────────────────────────────────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────────┐ ┌──────────┐
│ EMERGENCY│ │ TRIAGE │ │ BOOKING │
│ GUARDRAIL│ │ RAG SEARCH │ │ HANDLER │
└──────────┘ └──────┬───────┘ └──────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────┐
│ RETRIEVAL PIPELINE │
│ │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────────────────┐ │
│ │ NVIDIA NIM │ │ BM25 │ │ CrossEncoder Reranker │ │
│ │ Embeddings │ + │ Keyword │ → │ ms-marco-MiniLM-L-6 │ │
│ │ (Semantic) │ │ Matching │ │ (Top-10 → Top-3) │ │
│ └──────┬──────┘ └──────┬──────┘ └───────────┬─────────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ PINECONE VECTOR DATABASE │ │
│ │ 90,306 vectors • cosine • 1024 dimensions │ │
│ └─────────────────────────────────────────────────────────────┘ │
└────────────────────────────┬────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────┐
│ SYNTHESIS (Groq llama-3.1-8b-instant) │
│ Grounded response + clinical justification + source attribution │
└────────────────────────────┬────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────┐
│ PATIENT RESPONSE │
└─────────────────────────────────────────────────────────────────────┘
Benchmark
| Metric | Value |
|---|---|
| Intent Accuracy | 96.0% |
| Avg Latency | 9.26s |
| Avg Retrieval Score | 0.92 |
| Records in DB | 90,306 |
| Intent | Correct | Total | Accuracy |
|---|---|---|---|
| EMERGENCY | 10 | 10 | 100% |
| TRIAGE | 18 | 20 | 90% |
| BOOKING | 10 | 10 | 100% |
| GENERAL_INFO | 5 | 5 | 100% |
| CLOSING | 5 | 5 | 100% |
Tech Stack
| Layer | Technology | Purpose |
|---|---|---|
| SDK | ella-sdk (PyPI) |
Python client |
| Embeddings | NVIDIA NIM (nv-embedqa-e5-v5) |
1024-dim semantic vectors |
| Vector DB | Pinecone (Serverless, AWS) | Cosine similarity search |
| LLM | Groq (llama-3.1-8b-instant) |
Intent classification + response generation |
| Reranker | CrossEncoder (ms-marco-MiniLM-L-6-v2) |
Precision reranking |
| Orchestration | LangChain + LangGraph | Agent pipeline |
| Validation | Pydantic | Schema-validated outputs |
License
MIT License — see LICENSE for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ella_sdk-1.1.7.tar.gz.
File metadata
- Download URL: ella_sdk-1.1.7.tar.gz
- Upload date:
- Size: 28.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1daeac6bd061553324cf627187765b6a9babfe3d5d9ac99bbfe049575ec2b2b2
|
|
| MD5 |
7042ea9a5f7fd8fdd011a2c908d9412d
|
|
| BLAKE2b-256 |
31c3463e4c7ed94f60428f90aa1dccb7723810a59ad8f815fd93485549b1367a
|
File details
Details for the file ella_sdk-1.1.7-py3-none-any.whl.
File metadata
- Download URL: ella_sdk-1.1.7-py3-none-any.whl
- Upload date:
- Size: 36.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b8794312413db7ae4875d4f640ecfff79207a794f50d9d51a2dc03c347995d0d
|
|
| MD5 |
b4a06ee3d5ade48ccd89d5bbe2d6d273
|
|
| BLAKE2b-256 |
77ab2ecfc92e1a4dc21d8f367b8bbb3ca1499ffa95cf272814c4777b5558f30b
|