Skip to main content

Python SDK for Om

Project description

OMTX Python SDK

Minimal Python client for the OMTX Gateway.

Core scope:

  • Submit diligence jobs
  • Poll job status/results
  • Request subscription-gated shard URLs
  • Load entitlement-scoped data into an OmDataFrame (Polars-backed)

Installation

Core SDK:

pip install omtx

Quick Start

from omtx import OmClient

with OmClient() as client:
    print(client.status())

    job = client.diligence.deep_diligence(
        query="CRISPR applications in cancer therapy",
        preset="quick",
    )

    result = client.jobs.wait(
        job["job_id"],
        result_endpoint="/v2/jobs/deep-diligence/{job_id}",
    )
    print(result.get("result", {}).get("total_claims"))

Setup

export OMTX_API_KEY="your-api-key"

The SDK targets https://api.omtx.ai.

For develop/staging testing, override the base URL explicitly:

gcloud run services describe api-gateway \
  --region us-central1 \
  --project omtx-diligence \
  --format=json | jq -r '.status.traffic[]? | select(.tag=="develop") | .url'

Then use that URL in the client:

from omtx import OmClient

with OmClient(
    api_key="your-api-key",
    base_url="https://develop---api-gateway-...a.run.app",
) as client:
    print(client.status())

Or with env:

export OMTX_BASE_URL="https://develop---api-gateway-...a.run.app"

When a non-production host is used, the SDK emits a warning.

Data Access

Primary training flow (separate pools):

binders = client.load_binders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=1000,          # optional: random sample size
    sample_seed=42,  # optional: deterministic sampling
)

nonbinders = client.load_nonbinders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=10000,         # optional: random sample size
    sample_seed=42,  # optional: deterministic sampling
)

print(binders.shape, nonbinders.shape)
binders.show(top_n=24, sort_by="binding_score")
binders.show(top_n=24, sort_by="selectivity_score")

Manual shard export URLs (advanced use):

urls = client.binders.urls(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
)
print("Binder shard URLs:", len(urls["binder_urls"]))
print("Non-binder shard URLs:", len(urls["non_binder_urls"]))
print("First binder URL:", urls["binder_urls"][0] if urls["binder_urls"] else None)

Generated proteins available now:

protein_uuids = client.datasets.generated_protein_uuids()
print("Generated protein UUIDs:", protein_uuids[:5])

Module-level convenience:

import omtx as om

binders = om.load_binders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=1000,
    sample_seed=42,
)
nonbinders = om.load_nonbinders(
    protein_uuid="550e8400-e29b-41d4-a716-446655440000",
    n=10000,
    sample_seed=42,
)
print(binders.shape, nonbinders.shape)

Chemprop Training (Binary Binder Classification)

Use load_binders(...) and load_nonbinders(...) to build a labeled dataset (is_binder=1/0) and train Chemprop in classification mode.

Prepare training CSV:

import polars as pl
from omtx import OmClient

PROTEIN_UUID = "550e8400-e29b-41d4-a716-446655440000"

with OmClient() as client:
    binders_df = (
        client.load_binders(
            protein_uuid=PROTEIN_UUID,
            n=50000,
            sample_seed=42,
        )
        .to_polars()
        .with_columns(pl.lit(1).alias("is_binder"))
    )
    non_binders_df = (
        client.load_nonbinders(
            protein_uuid=PROTEIN_UUID,
            n=200000,
            sample_seed=42,
        )
        .to_polars()
        .with_columns(pl.lit(0).alias("is_binder"))
    )

train_df = (
    pl.concat([binders_df, non_binders_df], how="vertical_relaxed")
    .select(["smiles", "is_binder"])
    .drop_nulls()
    .unique()
    .sample(fraction=1.0, shuffle=True)
)

train_df.write_csv("chemprop_train.csv")
print(train_df.shape)

Equivalent script:

python examples/prepare_chemprop_binary.py \
  --protein-uuid 550e8400-e29b-41d4-a716-446655440000 \
  --binders 50000 \
  --non-binders 200000 \
  --output chemprop_train.csv

Train:

chemprop train \
  --data-path chemprop_train.csv \
  --task-type classification \
  --smiles-columns smiles \
  --target-columns is_binder \
  --split-type scaffold_balanced \
  --split-sizes 0.8 0.1 0.1 \
  --epochs 50 \
  --batch-size 64 \
  --output-dir chemprop_runs/binder_cls

Predict:

chemprop predict \
  --test-path infer.csv \
  --model-paths chemprop_runs/binder_cls \
  --smiles-columns smiles \
  --preds-path infer_preds.csv

Idempotency

  • Every non-GET call gets an idempotency key automatically.
  • All diligence POST helpers accept idempotency_key=... for explicit retry control.

Helper Surface

  • diligence.deep_diligence(query, preset=None, idempotency_key=None, **kwargs)
  • diligence.synthesize_report(gene_key, idempotency_key=None)
  • diligence.search(query, idempotency_key=None)
  • diligence.gather(query, idempotency_key=None)
  • diligence.crawl(url, max_pages=5, idempotency_key=None)
  • diligence.list_gene_keys()
  • jobs.history(...), jobs.status(job_id), jobs.wait(job_id, ...)
  • binders.get_shards(...)
  • binders.urls(...)
  • load_binders(...)
  • load_nonbinders(...)
  • datasets.catalog()
  • datasets.generated_protein_uuids()
  • status()
  • users.profile()

Route policy:

  • /v2/diligence/getTargetDiligenceReport remains an alias route and is not a separate SDK helper.
  • /v2/rag/search is intentionally not exposed in the SDK.

Migration

Breaking changes in 2.0.0:

  • OMTXClient removed.
  • OmClient is now the only supported client class.
  • Legacy pricing helpers removed from SDK surface.
  • Legacy binder batch-cost helper removed from SDK surface.
  • Shard access now resolves latest accessible dataset by protein_uuid.
  • client.status() is the primary health helper.
  • load_binders(...) and load_nonbinders(...) are the primary dataframe-loading helpers.
  • Flat shard URL aliases are available as binder_urls / non_binder_urls.
  • Core SDK runtime includes polars + rdkit.

Migration mapping (1.x -> 2.x):

  • from omtx import OMTXClient -> from omtx import OmClient
  • OMTXClient(...) -> OmClient(...)

Breaking changes in 1.0.0:

  • binders.get(...) removed from core SDK.
  • binders.iter(...) removed from core SDK.
  • pandas removed from required dependencies.

Migration mapping (0.x -> 1.x):

  • binders.get(...) -> client.load_binders(...) / client.load_nonbinders(...) or binders.get_shards(...)
  • binders.iter(...) -> binders.get_shards(...) + application-level streaming
  • pip install omtx (with pandas) -> pip install omtx (with polars + rdkit)

Full details: see MIGRATION.md.

Requirements

  • Python >=3.9
  • OMTX API key

License

MIT. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

omtx-2.0.1.tar.gz (22.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

omtx-2.0.1-py3-none-any.whl (18.3 kB view details)

Uploaded Python 3

File details

Details for the file omtx-2.0.1.tar.gz.

File metadata

  • Download URL: omtx-2.0.1.tar.gz
  • Upload date:
  • Size: 22.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.14

File hashes

Hashes for omtx-2.0.1.tar.gz
Algorithm Hash digest
SHA256 72ed28f34015f7340a7aa046d5802d264d32f4c7a726d7c15052b1584ef52acb
MD5 0fbfcf8d8154d1a2f875993e0d28cde7
BLAKE2b-256 fba2ce0efab6ee694ba2a225cd410f40f2c4a2b7589b8a5b7a78013468a47832

See more details on using hashes here.

File details

Details for the file omtx-2.0.1-py3-none-any.whl.

File metadata

  • Download URL: omtx-2.0.1-py3-none-any.whl
  • Upload date:
  • Size: 18.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.14

File hashes

Hashes for omtx-2.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 45a60153252d498a4d020a5a66fce93a0286466784887f7f36c53189da4668d3
MD5 dd16de582f21bdf3a7e0cb7ec025fb10
BLAKE2b-256 23f43f61f115da8fe6234bfbfd841163deaed7a5ee9ea27d91613064af98881a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page