Skip to main content

OnToma

Introduction

OnToma is a Python package for mapping entities to identifiers using lookup tables. It is optimised for large-scale entity mapping, and is designed to work with PySpark DataFrames.

OnToma supports the mapping of two kinds of entities: labels (e.g. brachydactyly) and ids (e.g. OMIM:112500).

OnToma includes a NER (Named Entity Recognition) module for extracting clean entity names from raw text labels. This is useful when your data contains labels that need preprocessing. Currently, this feature is available for drugs and diseases. To use NER features, see NER Module Documentation.

OnToma currently has modules to generate lookup tables from the following datasources:

  • Open Targets disease, target, and drug indices
  • Disease curation tables with the SEMANTIC_TAG and PROPERTY_VALUE fields (e.g. the Open Targets disease curation table)
  • You can also provide your own curation tables as long as they are compatible with the defined schema

The package features entity normalisation using Spark NLP, where entities in both the lookup table and the input dataframe are normalised to improve entity matching.

Successfully mapped entities may be mapped to multiple identifiers.

Prerequisites

Java Runtime Environment

OnToma requires a Java runtime, as it's a prerequisite for PySpark and Spark-NLP: OpenJDK 8, 11 or 17 for Spark 3.5, and OpenJDK 17 or 21 for Spark 4. OpenJDK 17 works with both.

macOS Installation

Install OpenJDK 17 using Homebrew:

brew install openjdk@17

After installation, you need to set the JAVA_HOME environment variable. Add the following to your shell configuration file (e.g., ~/.zshrc or ~/.bash_profile):

export JAVA_HOME="/opt/homebrew/opt/openjdk@17/libexec/openjdk.jdk/Contents/Home"
export PATH="$JAVA_HOME/bin:$PATH"

Reload your shell configuration:

source ~/.zshrc

Verify the installation:

java -version

Installation

pip install ontoma

Spark session configuration

OnToma requires a Spark session configured to include the Spark NLP library. OnToma supports Spark 3.5 and Spark 4; Spark 4 needs Spark NLP 7.0.0 or newer. spark_nlp_coordinate() returns the Spark NLP artifact matching the installed PySpark and Spark NLP versions (Scala 2.12 for Spark 3, Scala 2.13 for Spark 4).

from pyspark.sql import SparkSession
from pyspark.conf import SparkConf

from ontoma import spark_nlp_coordinate

# add Spark NLP library to Spark configuration
config = (
    SparkConf()
    .set("spark.jars.packages", spark_nlp_coordinate())
)

# create Spark session
spark = SparkSession.builder.config(conf=config).getOrCreate()

Usage example

Here is an example showing how OnToma can be used to map diseases:

First, load data to generate a disease label lookup table:

from ontoma import OnToma, OpenTargetsDisease

disease_index = spark.read.parquet("path/to/disease/index")
disease_label_lut = OpenTargetsDisease.as_label_lut(disease_index)

Then, create the OnToma object to be used for mapping entities:

ont = OnToma(
    spark = spark, 
    entity_lut_list = [disease_label_lut]
)

Given an input PySpark DataFrame disease_df containing the diseases to be mapped in the column disease_name:

mapped_disease_df = ont.map_entities(
    df = disease_df,
    result_col_name = "mapped_ids",
    entity_col_name = "disease_name",
    entity_kind = "label",
    type_col = f.lit("DS")
)

Mapping results can be found in the column mapped_ids. The results will be in the form of a list of identifiers that the entity is successfully mapped to.

Using NER for preprocessing (drugs)

When your drug labels contain dosages, forms, or brand names, use the NER module to extract clean entity names before mapping:

from ontoma.ner.drug import extract_drug_entities
import pyspark.sql.functions as f

# Extract clean drug entities from raw labels
df_extracted = extract_drug_entities(
    spark=spark,
    df=raw_drug_df,
    input_col="raw_drug_label",
    output_col="extracted_drugs"
)

# Explode arrays for mapping
df_exploded = df_extracted.select("*", f.explode("extracted_drugs").alias("clean_drug"))

# Map with OnToma
mapped_df = ont.map_entities(
    df=df_exploded,
    entity_col_name="clean_drug",
    entity_kind="label",
    type_col=f.lit("drug")
)

See NER Module Documentation for more details.

Speeding up subsequent OnToma usage

PySpark uses lazy evaluation, meaning transformations are not executed until an action is triggered.

When using the same OnToma object multiple times, it is recommended to specify a cache directory when creating the OnToma object using the cache_dir parameter to avoid re-running the lookup table processing logic on each use.

ont = OnToma(
    spark = spark, 
    entity_lut_list = [disease_label_lut],
    cache_dir = "path/to/cache/dir"
)

Development

Running Tests

Install development dependencies:

uv sync --dev

Run all tests:

uv run pytest

Skip slow tests (e.g., NER tests that download large models):

uv run pytest -m "not slow"

Metadata

Release files for ontoma 2.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ontoma 2.7.0
File Size Uploaded
ontoma-2.7.0.tar.gz 29.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ontoma 2.7.0
File Interpreter ABI Platform
ontoma-2.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 71.0 kB

Release files / ontoma-2.7.0.tar.gz

Download URL ontoma-2.7.0.tar.gz
Size 29.4 kB
Tags Source
SHA-256 checksum
How to use checksums
703a92abba13b84decc97f62e039e362a66cfd9e0f2deaa641ca754fa1607d5f
BLAKE2b-256 checksum
How to use checksums
b4c3e8eafcb69720355c9ba1198593e1154471176710ab3bd544be657181250c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / ontoma-2.7.0-py3-none-any.whl

Download URL ontoma-2.7.0-py3-none-any.whl
Size 41.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b2b80fdfc381a747b133041ff5717886508c7e657e5ce55a1db3d96b1e2bd9c1
BLAKE2b-256 checksum
How to use checksums
6991ec4a7f9a18248f598fffb5a22d6c0e8d8f9ffabf0fe2a371823882052e0f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

2.7.1

2 release files

This release

2.7.0 This release

2 release files

2.6.1

2 release files

2.6.0

2 release files

2.5.3

2 release files

2.5.1

2 release files

2.5.0

2 release files

2.4.2

2 release files

2.4.1

2 release files

2.4.0

2 release files

2.3.1

2 release files

2.3.0

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.0

2 release files

1.1.2

2 release files

1.1.0

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.0.18

2 release files

0.0.17

1 release file

0.0.16

1 release file

0.0.15

1 release file

0.0.14

1 release file

0.0.13

2 release files

0.0.11

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page