Skip to main content

Semantic search, PII masking, and schema understanding for Polars DataFrames

Project description

Omna

PyPI Python License Tests

Hybrid search (semantic + keyword), enterprise-grade PII detection & masking, and schema understanding — directly on your Polars DataFrames. No vector database. No API key. Data never leaves your machine.


The problem

# Finding every insurance claim denial — painful
keywords = ["claim denied", "coverage rejected", "policy voided", ...]
pattern  = re.compile("|".join(keywords), re.IGNORECASE)
results  = df[df["text"].str.contains(pattern, na=False)]
# Still misses: "insurer refused to honour the policy"
# Still misses: "claim outcome: not payable"
# Still misses: medical claim rejections using clinical terminology
# ...50+ lines per task. Grows with every edge case. Still wrong.
# With Omna
results = df.omna.search("insurance claim denied", on="text", k=5)
# Finds ALL of them — including docs that never say "denied" literally.
# 9ms. 50,000 documents. Zero cloud.

filtered = df.omna.filter("insurance claim denied", on="text", threshold=0.73)
# Every semantically matching document above the threshold.
# No keyword lists. No guesswork. Pure meaning.

answer = results.omna.ask("What personal data do these documents expose?")
# → "These insurance documents expose SSNs, medical record numbers,
#    dates of birth, health plan numbers, and claimant identifiers."
# Instant. One line.

# Auditing for PII before the data ships — painful
for col in df.columns:
    for i, val in enumerate(df[col].to_list()):
        if re.search(r'\b\d{3}-\d{2}-\d{4}\b', str(val)):   # SSNs only
            print(f"row {i}, {col}: {str(val)[:60]}")
# Catches one pattern. Misses emails, phones, names, IBANs.
# No confidence score. No audit trail. No redaction.
# With Omna
df.omna.pii_report()          # audit — find every leak, every column
df.omna.mask_pii()            # redact — one line, full audit log
df.omna.mask_pii(model=True)  # + on-device AI model for contextual PII
# A six-layer engine: regex + checksum-validated IDs + 220+ secret rules +
# an on-device AI model — the SAME Rust engine as the Omna Mac app & extension.
# Reversible [PERSON_1] tokens; secrets always irreversibly redacted. Local.
# Benchmarked openly on Gretel (core-PII recall 0.84 with the model, up from
# 0.69 on the prior engine): see docs/benchmark.md.

Demo

The Sword — hybrid search (semantic + keyword), filter, and ask across 50,000 documents:

Omna Sword Demo

The Shield — PII audit and redaction in one line:

Omna Shield Demo

Dataset: Gretel PII Benchmark (acquired by NVIDIA) — 50,000 synthetic documents built to test data privacy tools.


Enterprise-grade PII masking

Most "PII for DataFrames" tools are a regex or a Presidio wrapper. Omna's masking is a six-layer detection engine (pure Rust, fully on-device) — the same engine that powers the Omna Mac app and browser extension:

  • L1 — patterns + checksum validators: emails, SSNs, cards (Luhn), IBANs, and 30+ international IDs verified, not just pattern-matched
  • L2 — secrets: 220+ rules (AWS keys, GitHub tokens, JWTs, private keys) with entropy checks
  • L3 — on-device AI model (mask_pii(model=True)): catches contextual PII no regex can — bare names, addresses, medical context
  • L4–L6: entity resolution, reversible [PERSON_1] tokens (secrets always irreversibly redacted), and a full audit trail

No cloud, no API key, no data leaves your machine.

Same Gretel benchmark, same scoring — only the engine changed:

Gretel benchmark Before (Presidio) After (Omna engine)
Core-PII recall 0.69 0.84
All-types recall 0.35 0.79
All-types F1 0.50 0.82
Types · secret rules · validated IDs ~17 · 0 · none 30+ · 220+ · 30+
df.omna.pii_report()          # audit — every PII column, with confidence
df.omna.mask_pii()            # redact (L1+L2) — instant, full audit log
df.omna.mask_pii(model=True)  # full L1–L6, model-grade

Full methodology + numbers: docs/benchmark.md.


Install

pip install "omna[all]"

Supported platforms: Python 3.10–3.12 on macOS or Linux. Windows is not supported. The PII engine (omna[pii], and mask_pii(model=True)) runs on Apple Silicon macOS and Linux — it is not available on Intel Macs, because its on-device model runtime (ONNX Runtime) ships no Intel-Mac build; Intel-Mac users can still use search, filter, and schema understanding.

Extras: omna[embed] (search/filter), omna[pii] (PII detection & masking), omna[ask] (LLM queries). A bare pip install omna (just polars + numpy + rich) gives schema understanding (understand_df()); the heavier features are opt-in. PII masking is powered by the compiled omna-pii-mask engine — a self-contained wheel with no heavy Python ML dependencies, pulled in automatically by omna[pii], omna[ask] (which masks rows before the API call), or omna[all]. No API key needed for search, filter, embed, pii_report, mask_pii, or understand. Only ask() requires ANTHROPIC_API_KEY.


Quick start

import polars as pl
import omna

df = pl.read_csv("documents.csv")

# 1 — explore the schema
omna.understand_df(df)

# 2 — audit for PII before anything touches the data
df.omna.pii_report()

# 3 — redact
clean = df.omna.mask_pii()

# 4 — build a search index once
clean.omna.embed("text")

# 5 — search by meaning
results = clean.omna.search("insurance claim denied", on="text", k=5)

# 6 — filter everything above a threshold
flagged = clean.omna.filter("insurance claim denied", on="text", threshold=0.73)

# 7 — ask a question in plain English
results.omna.ask("What personal data do these documents expose?")

What Omna does

Method What it does
omna.understand_df(df) Schema inference — labels, null rates, samples. No LLM.
df.omna.embed(column) Vectorize a text column once; reuse across sessions
df.omna.search(query, on, k) Top-k results by hybrid relevance (semantic + keyword)
df.omna.filter(query, on, threshold) Every row above a similarity threshold
df.omna.pii_report() Audit every string column for PII
df.omna.mask_pii() Redact PII, auto-save audit log
df.omna.ask(question) Natural language queries over your DataFrame

API reference

omna.understand_df(df) — explore before you do anything

No LLM. No API call. Analyzes column names, dtypes, null rates, and sample values.

omna.understand_df(df)
 column                dtype    null_pct   label     sample
 uid                   String     0.0%     category  24bb757...
 domain                String     0.0%     category  insurance, healthcare...
 document_type         String     0.0%     category  Invoice, ClaimForm...
 document_description  String     0.0%     text      An insurance claim...
 text                  String     0.0%     text      **Claim ID: 285-14...

Labels: email phone name id date text numeric boolean category unknown

df.omna.embed(column) — vectorize once, search forever

Converts text to 768-dimensional vectors using FastEmbed (local ONNX, no API key). Saves to .omna/{column}.parquet. Run once — search() and filter() load it automatically on every subsequent call.

df.omna.embed("text")
# → .omna/text.parquet

Model: nomic-ai/nomic-embed-text-v1.5 (768-dim, downloaded once on first use) — the same embedding model the Omna Mac app uses. Embed is a one-time cost.

Hardware 50k rows
MacBook Air M5 ~45 min
MacBook Pro M4 Max ~15 min
AWS GPU instance ~2 min
df.omna.search(query, on, k) — hybrid search (semantic + keyword)

Requires df.omna.embed("column") first.

results = df.omna.search("insurance claim denied", on="text", k=5)

Search is hybrid by default: semantic (embedding) similarity and BM25 keyword matching run together, fused with Reciprocal Rank Fusion. Semantics catch meaning; BM25 catches rare exact tokens the embeddings blur — part codes, IDs, surnames, acronyms. Results are ordered by fused relevance. The BM25 index is built from the saved column on first use (no re-embedding, no change to your index file). Pass hybrid=False for pure-semantic search.

results = df.omna.search("XJ9000", on="parts", k=5)              # exact code → BM25 nails it
results = df.omna.search("claim denied", on="text", hybrid=False) # meaning only
 uid            document_type         domain      text                               _score
 67fccc1e207…   ClaimSummary          insurance   **Claim ID: 285-14-1755, Policy…   0.762
 b8ae088cd21…   ClaimSummary          insurance   **Claim Summary**…                 0.749
 de5bba0a2cc…   Insurance Claim Form  healthcare  **Insurance Claim Form**…          0.748
 ebccdde3b42…   Insurance Claim       healthcare  Insurance Claim for MED74974358…   0.747
 aebb0eb55fb…   ClaimForm             healthcare  **Claim Form** - Patient ID…       0.747

_score is cosine similarity (0–1). None of these documents contain the phrase "insurance claim denied" — Omna finds them by meaning.

df.omna.filter(query, on, threshold) — semantic filter

Requires df.omna.embed("column") first.

filtered = df.omna.filter("insurance claim denied", on="text", threshold=0.73)
# → N documents matched — all semantically related to claim denials

Returns every row above the threshold. Default: 0.3. Raise for precision, lower for recall.

Use search() for the top k. Use filter() for everything above a threshold.

df.omna.pii_report() — audit before you redact
df.omna.pii_report()
 column    detected types                                    hit rate   flagged
 entities  CREDIT_CARD, EMAIL_ADDRESS, PERSON, PHONE_NUMBER   85.4%    ✓ YES
 text      CREDIT_CARD, EMAIL_ADDRESS, PERSON, PHONE_NUMBER   78.1%    ✓ YES

Scans every string column. Returns hit rates, PII types, and confidence scores. Nothing is modified.

df.omna.mask_pii() — redact in one line
clean = df.omna.mask_pii()
# → Omna's own six-layer Rust engine (same kernel as the Mac app + browser
#   extension): reversible [PERSON_1]-style tokens; credentials/secrets
#   always irreversibly [REDACTED:KIND]; checksum-validated IDs; 220+ secret
#   rules. No heavy Python ML dependencies.
# → audit log saved to .omna/pii_audit.parquet automatically

clean = df.omna.mask_pii(model=True)
# → adds L3, the on-device AI model, for contextual PII regex can't catch
#   (bare prose names, addresses). Downloads the model (~809 MB) once.
# Requires the omna_pii_mask wheel (built from omna-workspace); see CHANGELOG.md.

Detects: PERSON EMAIL PHONE CREDIT_CARD US_SSN IP_ADDRESS IBAN MEDICAL_RECORD_NUMBER BANK_ACCOUNT, 220+ secret types, and 30+ international IDs — checksum-validated where applicable.

# Add the on-device AI layer (L3) for contextual PII regex can't catch —
# bare names, addresses, medical context. Downloads the model once.
clean = df.omna.mask_pii()              # L1+L2 (instant, deterministic)
df.omna.ask(question) — natural language queries

Privacy (since 2026-06-10): the sampled rows in the prompt are masked before they leave your machine (PII replaced with tokens; secrets redacted). Pass mask_rows=False to send raw rows for synthetic/public data.

Sends schema + up to 20 sample rows to Claude. Requires ANTHROPIC_API_KEY.

export ANTHROPIC_API_KEY=sk-ant-...
results.omna.ask("What personal data do these documents expose?")
# → "These insurance documents expose SSNs, medical record numbers,
#    dates of birth, health plan numbers, and claimant identifiers."

# Override model
results.omna.ask("Summarise the key themes", model="claude-sonnet-4-6")

Default model: claude-haiku-4-5-20251001.


How it works

df.omna.search("insurance claim denied", on="text", k=5)
         │
         ▼
   embedder.py       FastEmbed — nomic-embed-text-v1.5, local ONNX
                     query → [0.12, -0.34, 0.87, ...]  768-dim vector
         │
         ▼
   index.py          loads .omna/text.parquet → Arrow memory, zero-copy
                     50,000 stored vectors in Polars' own allocation
         │
         ▼
   similarity.rs     Rust kernel — cosine similarity over all vectors
                     returns top-k sorted descending, no Python loop
         │
         ▼
   frame.py          slices result rows, attaches _score → pl.DataFrame

The Rust kernel is under 70 lines. Dot products and norms in machine code, no intermediate allocations. 500,000 × 768-dim vectors scored in milliseconds on a single core.


Performance

50k rows 500k rows
Omna search 9ms 27ms
Omna filter 9ms 27ms
Pandas + FAISS ~25ms + index build ~25ms + index build
Polars keyword regex 1ms — exact match only 1ms — exact match only

Benchmarked on MacBook Air M5 with the prior 384-dim model, 10-query median, warm index. The current model (nomic-embed-text-v1.5, 768-dim) roughly doubles the per-query cosine cost — still single-digit-to-low-tens of milliseconds at these sizes.

Omna inherits Polars' Arrow columnar memory. The Rust similarity kernel operates on the same memory — no copy into NumPy, no copy into a C buffer.


FAQ

Does Omna send my data to the cloud?

No. Embedding, search, filter, PII detection, and masking all run locally. The only method that makes a network call is ask(), which sends schema metadata and sample rows to Claude via the Anthropic API — and only when you explicitly call it.

Do I need a GPU?

No. FastEmbed uses ONNX and runs on CPU. On Apple Silicon, it uses CoreML automatically. Embedding 50,000 documents takes ~45 minutes on a MacBook Air M5 — a one-time cost. After that, search() and filter() run in milliseconds from the saved index.

Why not FAISS / ChromaDB / Pinecone?

Those are vector databases. Omna is a Polars plugin. If your data already lives in a DataFrame, Omna adds hybrid search (semantic + keyword) with zero infrastructure — no separate process, no index server, no network hop. It's the difference between df.omna.search(...) and spinning up a separate service just to query your own data.

What PII types does Omna detect?

PERSON, EMAIL, PHONE, CREDIT_CARD, US_SSN, IP_ADDRESS, IBAN, MEDICAL_RECORD_NUMBER, BANK_ACCOUNT, 220+ secret types (API keys, tokens), and 30+ international IDs, and more. Detection runs on Omna's own six-layer Rust engine — fully local, no heavy Python ML dependencies.

Which Polars versions are supported?

Omna is tested on Polars 1.0+. It installs as a namespace plugin via df.omna.* — no import needed after import omna.

The embed step took 45 minutes. Do I have to redo it every time?

No. embed() saves the index to .omna/{column}.parquet. Every subsequent search() or filter() call loads it in ~300ms. You only re-run embed() if your data changes.


Roadmap

# Coming in v0.2
matched = transactions.omna.join(regulatory_categories, on="description")
# Match rows between two DataFrames by meaning, not exact key.

Star the repo to follow progress.


What's new

The PII engine was rebuilt from a Presidio + spaCy wrapper into Omna's own six-layer Rust engine (recall up ~2×, see Enterprise-grade PII masking), and the embedding model was upgraded to nomic-embed-text-v1.5. Full history in CHANGELOG.md.


License

Layer License
Python package (omna/) MIT
Rust engine (src/) Proprietary — ships as a compiled binary in the pip wheel

omna.dev · PyPI · GitHub

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

omna-0.2.1-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (286.8 kB view details)

Uploaded CPython 3.12manylinux: glibc 2.17+ x86-64

omna-0.2.1-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (271.6 kB view details)

Uploaded CPython 3.12manylinux: glibc 2.17+ ARM64

omna-0.2.1-cp312-cp312-macosx_11_0_arm64.whl (249.4 kB view details)

Uploaded CPython 3.12macOS 11.0+ ARM64

omna-0.2.1-cp312-cp312-macosx_10_12_x86_64.whl (257.1 kB view details)

Uploaded CPython 3.12macOS 10.12+ x86-64

omna-0.2.1-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (287.1 kB view details)

Uploaded CPython 3.11manylinux: glibc 2.17+ x86-64

omna-0.2.1-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (271.9 kB view details)

Uploaded CPython 3.11manylinux: glibc 2.17+ ARM64

omna-0.2.1-cp311-cp311-macosx_11_0_arm64.whl (251.1 kB view details)

Uploaded CPython 3.11macOS 11.0+ ARM64

omna-0.2.1-cp311-cp311-macosx_10_12_x86_64.whl (260.0 kB view details)

Uploaded CPython 3.11macOS 10.12+ x86-64

omna-0.2.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (287.4 kB view details)

Uploaded CPython 3.10manylinux: glibc 2.17+ x86-64

omna-0.2.1-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (272.1 kB view details)

Uploaded CPython 3.10manylinux: glibc 2.17+ ARM64

omna-0.2.1-cp310-cp310-macosx_11_0_arm64.whl (251.2 kB view details)

Uploaded CPython 3.10macOS 11.0+ ARM64

omna-0.2.1-cp310-cp310-macosx_10_12_x86_64.whl (260.0 kB view details)

Uploaded CPython 3.10macOS 10.12+ x86-64

File details

Details for the file omna-0.2.1-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 bbdddc413831bc7f5d39fafd64dfc2f2b080623cb284eb303903423a312cbca0
MD5 2c720a96a5b2ed72d077c8ab6788b25b
BLAKE2b-256 1fec518101416ce4f0c16e32f204937989731ea4baccb5acf0508b32a91deeaf

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 6e9a19edda101c98124e64f51a97b3aa504c6e78b63b82aa9b938e59fde37511
MD5 c9addd97559e75e0b164331fbe39e78c
BLAKE2b-256 5af7f64166283986826d0dc6f11ab8a74f6d8f0e88f24813aa57cdc77a40968e

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp312-cp312-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp312-cp312-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 e2f2932e5bc7709d302516c12a7cff58128fc871a6d6566ba5cc72a03036653a
MD5 5b5cd352470f8ac5f03d2b8158f0c445
BLAKE2b-256 4ede82d25c49d85b6bbc2a436473524453d9ecef8a986638e16fa97fdc8cb959

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp312-cp312-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp312-cp312-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 5967d086169ee41be2191e46fc8755bf6643dfaed1e3d9d3b4a81a6589aa6598
MD5 906c5c3bfe8526ea86aa7db4039ac724
BLAKE2b-256 631bb0b24968610c6e3366ea17b5f8d9f8ff631b1093c50993264adbfde0bde8

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 1e8a8e86979dc707cf38f6aaaf2943e9347203862b82f938fa6d6e78311897a5
MD5 b22d77a0d49065957d8ec42a00662ae4
BLAKE2b-256 ae50a01b8209715cbf06cd6b6157c346483f79f7ff18d49deae12e4e9cf22595

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 2142b4cebdfe0c6f164c9cab29426fbd8393ed5d71d7124fcf5cd00ebafb7549
MD5 d5b6259db6ca024032077805d423bb90
BLAKE2b-256 a95e4eac7965e7745203d70a359e57aea130b67a03f70cf314dccb304ece7c47

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp311-cp311-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp311-cp311-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 66d08553fe991f5ebbf938e188a798437a411a316e980bb99821f3bfea27dcfc
MD5 9d52c7079c3514cc0d9bd0b0d9f60ac1
BLAKE2b-256 73d1f10fa5b0dfeda273286d29f3cf9fb5d6e7cf689e3b76a400085ab3f99489

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp311-cp311-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp311-cp311-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 f387ef6c682f977c3c0c0aaa4dc2b25c62ba9fbff6196a3b6cbcc40a41058129
MD5 9c666f2f229d852c006651c78cb18467
BLAKE2b-256 2166eaf02e4ee2b99b01378aa73eac62a3720735609a38fc7bff542db4cb36bd

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 8c2b120d7fd41a8e5736a2621932d45c27794a06609488f3aa1f17451781f2b3
MD5 5880fb5cb823e11d232d0363a53a4ec0
BLAKE2b-256 9c0c9d490e10810ec9f9003cf59203f5a96d1b1ed4e5835f96bf771884c88a3d

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 f258d7c935ac3b8dcd4ef03dd8d993eba55724ba45275296de9f393ba21ab376
MD5 35b7dff9204a69cfdafef19fdd014fcd
BLAKE2b-256 5c99435d73886f64e74069f1a73b425f5df815c6013e6cdde0af9fc0b01141b5

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp310-cp310-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp310-cp310-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 539ac1a8f6c207ce45d1264ae55f1dac989c1e7eb4f43c3c9fcbbaaa841d068e
MD5 87042a144c63b56f6e0f038c60ae3fe1
BLAKE2b-256 9132a2b72ab3ca5070d86681b6056e130553ed54485128d0b206be2d35d166a6

See more details on using hashes here.

File details

Details for the file omna-0.2.1-cp310-cp310-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for omna-0.2.1-cp310-cp310-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 ccc0967381279f34c53e6af3f0427da8c3dc676a9d15d61d8bb99fe917c1262b
MD5 65dbc51bc5971a06bad928db22837c4b
BLAKE2b-256 c6938b0c9d46324d2eba1db157dc3e305aa060b9f5b69f6e69863a1d980d5fd8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page