Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

ClingRounder

An offset-safe clinical text grounding toolkit for extracting medical concepts, resolving clinical context, linking terminology, and validating relation graphs. It is designed for Vietnamese and mixed Vietnamese-English text while keeping the reusable contracts language-neutral.

ClingRounder is a reusable Python package and research portfolio. Deterministic rules, optional local model adapters, terminology repositories, neutral evaluation, and data-mining workflows share typed interfaces. Historical competition code is retained as an optional benchmark plugin; it is not part of the default runtime or evaluation path.

The PyPI distribution is named clingrounder; the Python import namespace remains clingrounder.

Research software only. It is not a medical device and must not be used as the sole basis for clinical decisions.

What It Does

  • Extracts diseases, symptoms, drugs, laboratory tests/results, procedures, findings, anatomy, and structured medication attributes such as strength, route, frequency, and duration.
  • Preserves exact raw-text spans through normalization, model tokenization, and export.
  • Classifies present, negated, historical, family, possible, planned, conditional, and resolved context when the configured context provider has evidence for the label.
  • Retrieves and links type-compatible ICD-10, RxNorm, and local terminology concepts.
  • Extracts typed relations and rejects invalid medical graph edges.
  • Builds derived SQLite FTS5 terminology and knowledge-graph indexes from canonical JSONL; JSONL remains the source of truth and stale derived indexes are rejected.
  • Evaluates spans, assertions, linking, relations, runtime, and error slices independently of a task.
  • Mines licensed data into provenance-rich bronze, silver, gold, and challenge snapshots.

Design Principles

  1. Raw offsets are authoritative. Normalized text is for lookup; exported spans always address the original source string.
  2. Terminology constrains linking. A code cannot be emitted unless it exists in the loaded, type-compatible terminology release.
  3. Composition is explicit. PipelineFactory is the composition root; PipelineRunner owns orchestration and receives concrete components through ports.
  4. Rules and models are replaceable. Deterministic baselines and local Hugging Face adapters implement the same contracts.
  5. Data and experiments are reproducible. Sources, configs, model revisions, fingerprints, prompts, and derived artifacts have explicit provenance.
  6. Benchmarks do not define the core. Task schemas, heuristics, exporters, and campaign records live below clingrounder.benchmarks.

Architecture

flowchart LR
    A[Raw document] --> B[Sections and sentences]
    B --> C[Entity proposal adapters]
    C --> D[Span and type resolution]
    D --> E[Assertion context graph]
    D --> F[Candidate retrieval]
    T[(JSONL / SQLite terminology)] --> C
    T --> F
    F --> G[Reranking and assignment]
    E --> H[Relations and KG checks]
    G --> H
    H --> I[Validated prediction]

    R[Rule adapters] --> C
    M[Local model adapters] --> C
    K[(SQLite knowledge graph)] --> G
    K --> H

The main dependency direction is:

schema + preprocessing + terminology ports
                    ↓
              pipeline ports
                    ↓
          rule and model adapters
                    ↓
             PipelineComponents
                    ↓
              PipelineRunner

generic evaluation ← task adapter ← optional benchmark plugin

See docs/architecture.md and docs/code-map.md for ownership and extension points.

Quickstart

Python 3.11 through 3.14 is supported.

git clone https://github.com/damminhtien/clingrounder.git
cd ontological-reasoning-in-medical-knowledge-retrieval

uv sync --extra dev
uv run clingrounder pipeline run \
  --config configs/pipeline/clinical-baseline.yaml \
  --input data/samples/sample_notes.jsonl \
  --output outputs/sample-predictions.jsonl

Without uv:

python -m venv .venv
source .venv/bin/activate
python -m pip install clingrounder
python -m pip install -e ".[dev]"
clingrounder pipeline run \
  --config configs/pipeline/clinical-baseline.yaml \
  --input data/samples/sample_notes.jsonl \
  --output outputs/sample-predictions.jsonl

The sample emits source-backed entities such as:

{
  "text": "viêm phổi",
  "span": [102, 111],
  "type": "DISEASE",
  "assertion": "POSSIBLE",
  "code_system": "ICD-10",
  "code": "J18.9"
}

Validate and evaluate the result:

uv run clingrounder validate \
  --profile development \
  --pred outputs/sample-predictions.jsonl \
  --documents data/samples/sample_notes.jsonl \
  --dictionary data/dictionaries/seed_concepts.jsonl

uv run clingrounder evaluate \
  --gold data/samples/gold.jsonl \
  --pred outputs/sample-predictions.jsonl \
  --error-analysis outputs/sample-errors.json

The installed CLI is split by responsibility: clingrounder exposes operational commands, clingrounder-research exposes mining/model commands, and clingrounder-benchmark loads optional benchmark plugins. They share one dispatcher and handler registry; no command implementation is duplicated. See docs/cli-scopes.md.

Python API

from clingrounder import Pipeline

with Pipeline.from_profile("clinical-baseline") as pipeline:
    prediction = pipeline.predict(
        "Bệnh nhân khó thở, không sốt.",
        document_id="note-001",
    )

for entity in prediction.entities:
    print(entity.text, entity.type.value, entity.assertion.value, entity.code)

The facade also provides predict_document, predict_many, and predict_with_trace. It owns terminology repositories, model adapters, caches, and worker resources and closes them when the context exits. Use Pipeline.from_config(path) for a checked-in or application-owned profile.

Advanced composition

Library and research integrations can compose the lower-level runtime explicitly:

from clingrounder.pipeline import PipelineComponents, PipelineFactory, PipelineRunner

components = PipelineComponents(...)  # inject ports and repositories explicitly
runner = PipelineRunner(components)
prediction = runner.process_text("note-001", "Bệnh nhân khó thở, không sốt.")

PipelineFactory remains the composition root for advanced integrations. Public ports include EntityExtractorPort, AssertionClassifierPort, CandidateRetrieverPort, CandidateRerankerPort, RelationExtractorPort, and TerminologyRepository; ordinary application code should use Pipeline instead.

Pipeline Profiles

Reusable profiles are explicit, path-stable YAML contracts:

Profile Purpose
configs/pipeline/clinical-baseline.yaml Small deterministic quickstart
configs/pipeline/full_terminology.yaml Full ICD-10/RxNorm normalization through SQLite
configs/pipeline/full_terminology_kg_exact.yaml Full terminology plus exact graph evidence
configs/pipeline/general_terminology_vn.yaml Experimental Vietnamese terminology profile
configs/pipeline/mined_vietbioner_silver.yaml Reviewed mined Vietnamese recognition overlay

clingrounder pipeline run has no hidden default profile. Model profiles must pin model_id and revision; model adapters are lazy and local-only by default. Install the ml extra only when using model-backed profiles.

Terminology At Scale

Canonical terminology remains JSONL. Runtime lookup uses a derived, content-addressed SQLite FTS5 index with read-only, query-only, thread-local connections.

uv run clingrounder terminology build \
  --source data/processed/full_concepts.jsonl \
  --cache-dir .cache/clingrounder/terminology

uv run clingrounder terminology inspect \
  --index .cache/clingrounder/terminology/<fingerprint>.sqlite3 \
  --query metformin \
  --entity-type DRUG \
  --code-system RxNorm

The index rejects stale source, schema, normalization, or alias fingerprints. Exact, abbreviation, lexical, BM25, optional dense, and graph-backed retrievers merge behind one retrieval pipeline; type and code-system filtering remains mandatory before assignment.

Research Portfolio

The repository includes several independently testable research tracks. Some are stable runtime components; model training, dense retrieval, graph evidence, and mining remain optional research workflows:

  • Proposal-first NER: dictionary, medication, lab, boundary, transformer, and generative adapters produce immutable evidence before global overlap resolution.
  • Structured medication linking: drug name, strength, administered dose, form, route, frequency, release, and brand are represented separately for RxNorm compatibility checks.
  • Context reasoning: assertion cues become a modifier-target graph with explicit scope, termination, priority, and provenance.
  • Hybrid retrieval: lexical and optional dense retrieval are separated from candidate qualification, reranking, and final assignment.
  • Graph evidence: exact linked concepts can provide bounded second-pass evidence without introducing new candidates.
  • Data mining: source connectors, immutable artifacts, parsers, deduplication, proposal labeling, review queues, coverage planning, and provenance-aware snapshots are reproducible stages. Access and license policy is checked before acquisition.

Start with docs/rule-ner.md, docs/reference-implementations.md, docs/data-mining.md, and docs/mining-reproducibility.md.

Data Mining And Provenance

uv run clingrounder-research data registry validate \
  --registry data/sources/mining_registry.yaml
uv run clingrounder-research data run --plan configs/mining/open_corpus_v1.yaml
uv run clingrounder-research data coverage report --help
uv run clingrounder-research data snapshot freeze --help

The public Git tree contains code, redistributable fixtures, policies, source dossiers, checksums, and rebuild instructions. Restricted clinical text, licensed terminology, manual labels, checkpoints, and generated runs remain in local or object storage. Their identities are recorded in data/provenance/local-artifacts.json and source-specific manifests.

Audit the publication boundary before release:

uv run clingrounder release audit \
  --policy configs/repository/public-release.yaml \
  --root .

See docs/public-release.md for restore and publication rules.

Optional Benchmark Plugin

The archived Vietnamese extraction challenge is retained for reproducibility and regression research. It is isolated from reusable pipeline defaults and has no stability guarantee:

uv run clingrounder-benchmark list
uv run clingrounder-benchmark phase1 --help
uv run pytest -o addopts='' -m "benchmark and not private and not model" \
  tests/benchmarks/phase1

Task configs are under configs/benchmarks/phase1. Restricted corpora and historical artifacts are restored by fingerprint and are not required for the toolkit quickstart.

Repository Map

src/clingrounder/
  pipeline/       ports, composition, runner, tracing, parallel batches
  ner/            proposal-first rules and structured span extractors
  adapters/       rule, hybrid, Hugging Face, and generative adapters
  context/        assertion cues, scope, and modifier graphs
  terminology/    repository contract and SQLite FTS5 backend
  retrieval/      retriever adapters and evidence fusion
  linking/        qualification, reranking, and assignment
  relations/      typed relation extraction
  kg/             graph storage, reasoning, and validation
  evaluation/     task-neutral metrics and reports
  mining/         source-to-snapshot data workflows
  benchmarks/     optional task plugins
configs/
  pipeline/       reusable runtime profiles
  mining/         source and curation plans
  benchmarks/     archived task profiles
tests/
  benchmarks/     opt-in benchmark suites

Development

# Fast unit and contract suite, normally under 15 seconds on the reference machine
uv run pytest tests

# All redistributable tests, including opt-in integration/release checks
uv run pytest -o addopts='' -m "not private and not model" tests

# Static checks
uv run ruff check .
uv run mypy src

Optional markers are integration, release, benchmark, private, and model. Tests touching schema, offsets, code systems, relation endpoints, or evidence spans remain hard gates.

Documentation

Licensed under the MIT License.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

clingrounder-0.1.0a1.tar.gz (1.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

clingrounder-0.1.0a1-py3-none-any.whl (1.1 MB view details)

Uploaded Python 3

File details

Details for the file clingrounder-0.1.0a1.tar.gz.

File metadata

  • Download URL: clingrounder-0.1.0a1.tar.gz
  • Upload date:
  • Size: 1.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for clingrounder-0.1.0a1.tar.gz
Algorithm Hash digest
SHA256 ea5cdc7868c55b3435b86a10b1494b61353438daedbbdcac3649a4c46f684836
MD5 a9d05395ea76581af10c86e5c3fda943
BLAKE2b-256 b2d0ca50e1b54dfd619cced104a1bb647d0ce247c2f109616b317920b8f55aa4

See more details on using hashes here.

Provenance

The following attestation bundles were made for clingrounder-0.1.0a1.tar.gz:

Publisher: release.yml on damminhtien/clingrounder

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file clingrounder-0.1.0a1-py3-none-any.whl.

File metadata

  • Download URL: clingrounder-0.1.0a1-py3-none-any.whl
  • Upload date:
  • Size: 1.1 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for clingrounder-0.1.0a1-py3-none-any.whl
Algorithm Hash digest
SHA256 215569d6b7650dcdd1360129e1f3bc863687260b0c37bbb1e2bd2d4bed1b9892
MD5 16a67c6acc38703e255f30226b304b89
BLAKE2b-256 d5f955df89cfcf85f7e61fdd80c34f96fdc8d634b5723865e75467600ba4a5bb

See more details on using hashes here.

Provenance

The following attestation bundles were made for clingrounder-0.1.0a1-py3-none-any.whl:

Publisher: release.yml on damminhtien/clingrounder

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page