Skip to main content

Python SDK for Om

Project description

OMTX Python SDK

Minimal Python client for the OMTX Gateway.

Core scope:

  • Submit diligence jobs
  • Poll job status/results
  • Request subscription-gated shard URLs
  • Load entitlement-scoped data into an OmDataFrame (Polars-backed)

Installation

Core SDK:

pip install omtx

Quick Start

from omtx import OmClient

with OmClient() as client:
    print(client.status())

    job = client.diligence.deep_diligence(
        query="CRISPR applications in cancer therapy",
        preset="quick",
    )

    result = client.jobs.wait(
        job["job_id"],
        result_endpoint="/v2/jobs/deep-diligence/{job_id}",
    )
    print(result.get("result", {}).get("total_claims"))

Setup

export OMTX_API_KEY="your-api-key"

The SDK targets https://api.omtx.ai.

For develop/staging testing, override the base URL explicitly:

gcloud run services describe api-gateway \
  --region us-central1 \
  --project omtx-diligence \
  --format=json | jq -r '.status.traffic[]? | select(.tag=="develop") | .url'

Then use that URL in the client:

from omtx import OmClient

with OmClient(
    api_key="your-api-key",
    base_url="https://develop---api-gateway-...a.run.app",
) as client:
    print(client.status())

Or with env:

export OMTX_BASE_URL="https://develop---api-gateway-...a.run.app"

When a non-production host is used, the SDK emits a warning.

Data Access

Primary training flow (separate pools):

binders = client.load_binders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=1000,          # optional: random sample size
    sample_seed=42,  # optional: deterministic sampling
)

nonbinders = client.load_nonbinders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=10000,         # optional: random sample size
    sample_seed=42,  # optional: deterministic sampling
)

print(binders.shape, nonbinders.shape)
binders.show(top_n=24, sort_by="binding_score")
binders.show(top_n=24, sort_by="selectivity_score")
# show() renders inline in notebooks; no extra display() wrapper needed.

Manual shard export URLs (advanced use):

urls = client.binders.urls(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
)
print("Binder shard URLs:", len(urls["binder_urls"]))
print("Non-binder shard URLs:", len(urls["non_binder_urls"]))
print("First binder URL:", urls["binder_urls"][0] if urls["binder_urls"] else None)

Generated proteins available now:

protein_uuids = client.datasets.generated_protein_uuids()
print("Generated protein UUIDs:", protein_uuids[:5])

Module-level convenience:

import omtx as om

binders = om.load_binders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=1000,
    sample_seed=42,
)
nonbinders = om.load_nonbinders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=10000,
    sample_seed=42,
)
print(binders.shape, nonbinders.shape)

Chemprop Training (Binary Binder Classification)

Use load_binders(...) and load_nonbinders(...) to build a labeled dataset (is_binder=1/0) and train Chemprop in classification mode.

Prepare training CSV:

import polars as pl
from omtx import OmClient

PROTEIN_UUID = "550e8400-e29b-41d4-a716-446655440000"

with OmClient() as client:
    binders_df = (
        client.load_binders(
            protein_uuid=PROTEIN_UUID,
            n=50000,
            sample_seed=42,
        )
        .to_polars()
        .with_columns(pl.lit(1).alias("is_binder"))
    )
    non_binders_df = (
        client.load_nonbinders(
            protein_uuid=PROTEIN_UUID,
            n=200000,
            sample_seed=42,
        )
        .to_polars()
        .with_columns(pl.lit(0).alias("is_binder"))
    )

train_df = (
    pl.concat([binders_df, non_binders_df], how="vertical_relaxed")
    .select(["smiles", "is_binder"])
    .drop_nulls()
    .unique()
    .sample(fraction=1.0, shuffle=True)
)

train_df.write_csv("chemprop_train.csv")
print(train_df.shape)

Equivalent script:

python examples/prepare_chemprop_binary.py \
  --protein-uuid 550e8400-e29b-41d4-a716-446655440000 \
  --binders 50000 \
  --non-binders 200000 \
  --output chemprop_train.csv

Train:

chemprop train \
  --data-path chemprop_train.csv \
  --task-type classification \
  --smiles-columns smiles \
  --target-columns is_binder \
  --split-type scaffold_balanced \
  --split-sizes 0.8 0.1 0.1 \
  --epochs 50 \
  --batch-size 64 \
  --output-dir chemprop_runs/binder_cls

Predict:

chemprop predict \
  --test-path infer.csv \
  --model-paths chemprop_runs/binder_cls \
  --smiles-columns smiles \
  --preds-path infer_preds.csv

Idempotency

  • Every non-GET call gets an idempotency key automatically.
  • All diligence POST helpers accept idempotency_key=... for explicit retry control.

Helper Surface

  • diligence.deep_diligence(query, preset=None, idempotency_key=None, **kwargs)
  • diligence.synthesize_report(gene_key, idempotency_key=None)
  • diligence.search(query, idempotency_key=None)
  • diligence.gather(query, idempotency_key=None)
  • diligence.crawl(url, max_pages=5, idempotency_key=None)
  • diligence.list_gene_keys()
  • jobs.history(...), jobs.status(job_id), jobs.wait(job_id, ...)
  • binders.get_shards(...)
  • binders.urls(...)
  • load_binders(...)
  • load_nonbinders(...)
  • datasets.catalog()
  • datasets.generated_protein_uuids()
  • status()
  • users.profile()

Route policy:

  • /v2/diligence/getTargetDiligenceReport remains an alias route and is not a separate SDK helper.
  • /v2/rag/search is intentionally not exposed in the SDK.

Migration

Breaking changes in 2.0.0:

  • OMTXClient removed.
  • OmClient is now the only supported client class.
  • Legacy pricing helpers removed from SDK surface.
  • Legacy binder batch-cost helper removed from SDK surface.
  • Shard access now resolves latest accessible dataset by protein_uuid.
  • client.status() is the primary health helper.
  • load_binders(...) and load_nonbinders(...) are the primary dataframe-loading helpers.
  • Flat shard URL aliases are available as binder_urls / non_binder_urls.
  • Core SDK runtime includes polars + rdkit.

Migration mapping (1.x -> 2.x):

  • from omtx import OMTXClient -> from omtx import OmClient
  • OMTXClient(...) -> OmClient(...)

Breaking changes in 1.0.0:

  • binders.get(...) removed from core SDK.
  • binders.iter(...) removed from core SDK.
  • pandas removed from required dependencies.

Migration mapping (0.x -> 1.x):

  • binders.get(...) -> client.load_binders(...) / client.load_nonbinders(...) or binders.get_shards(...)
  • binders.iter(...) -> binders.get_shards(...) + application-level streaming
  • pip install omtx (with pandas) -> pip install omtx (with polars + rdkit)

Full details: see MIGRATION.md.

Requirements

  • Python >=3.9
  • OMTX API key

License

MIT. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

omtx-2.0.2.tar.gz (22.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

omtx-2.0.2-py3-none-any.whl (18.5 kB view details)

Uploaded Python 3

File details

Details for the file omtx-2.0.2.tar.gz.

File metadata

  • Download URL: omtx-2.0.2.tar.gz
  • Upload date:
  • Size: 22.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.14

File hashes

Hashes for omtx-2.0.2.tar.gz
Algorithm Hash digest
SHA256 6045863cae3873fbd3664026670c16f3aa4a75c18e9169bb5d56c3b5eb4c5798
MD5 751eef846804f5d500f0af1341188b3b
BLAKE2b-256 7bb0e27ceec9c57023411838df48e6f647faf89f79924ab7c34429d0fe9f9dad

See more details on using hashes here.

File details

Details for the file omtx-2.0.2-py3-none-any.whl.

File metadata

  • Download URL: omtx-2.0.2-py3-none-any.whl
  • Upload date:
  • Size: 18.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.14

File hashes

Hashes for omtx-2.0.2-py3-none-any.whl
Algorithm Hash digest
SHA256 ab2505326ea5187e0d448079124487b888306b921399e032c7e647755d754067
MD5 2c12dab6200230d594e9d73677e3a2eb
BLAKE2b-256 3b0154b2fbe6ee97e3a25e0b0660a9b26a4eb84bd3d84da23371976edf200459

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page