Skip to main content

Universal data fetcher with Pydantic schemas

Project description

OmniFetcher

PyPI Version Python Versions CI License: MIT

One typed shape for every data source.

Jira, Postgres, PDFs, S3, Slack, Kafka, GitHub, Confluence — 30 sources, one contract. The code that reads one reads them all. Ships an MCP server, so your LLM agents can fetch from any of them through a single tool.

pip install omni-fetcher

The problem, in one screen

Pulling from four sources today means four SDKs, four auth models, four response shapes, and four ways to fail:

# GitHub — PyGithub
issue = Github(token).get_repo("psf/requests").get_issue(42)
text  = issue.body                                    # comments? separate call, separate shape

# Jira — atlassian-python-api
data  = Jira(url, username=u, password=t).issue("PROJ-1")
text  = data["fields"]["description"]                 # dict-diving, and it raises on 404

# A PDF — pypdf
text  = "\n".join(p.extract_text() for p in PdfReader("report.pdf").pages)

# Postgres — asyncpg
rows  = await (await asyncpg.connect(dsn)).fetch("SELECT id, email FROM users LIMIT 10")
text  = ...                                           # rows aren't text. and read-only? that's on you.

Now write the error handling. Four times. Now add a fifth source.

With OmniFetcher, it's one loop:

from omni_fetcher.v1 import AtomKind, OmniFetcher, Success, builtin_registry

omni = OmniFetcher(builtin_registry())

for uri in [
    "https://github.com/psf/requests/issues/42",
    "jira://issue/PROJ-1",
    "report.pdf",
    "postgres://db.internal/app?query=SELECT%20id,email%20FROM%20users%20LIMIT%2010",
]:
    result = await omni.fetch(uri, auth=creds_for(uri))   # your credential lookup
    if isinstance(result, Success):                       # -> one typed Result, every source
        text = "\n".join(a.content for a in result.tree.find_atoms(AtomKind.TEXT))
        print(uri, "→", len(text), "chars")

Same call. Same result type. Same three moves every time: check the state, walk the tree, read the atoms. No if source == "jira". Ever.

Who it's for

  • Building RAG or agent context — you need many heterogeneous sources in one shape, with provenance and honest failures, not thirty bespoke parsers.
  • Internal tools and glue services — one dependency instead of a dozen SDKs, and one error model instead of a dozen exception hierarchies.
  • Anyone wiring an LLM to real data — the MCP server does it with no glue code at all.

For agents: one MCP server, every source

pip install "omni-fetcher[mcp]"

omni-fetcher-mcp is a stdio MCP server that hands a model three tools:

Tool What it does
fetch(uri) The canonical result as JSON, for any bounded source
sample(uri, max_items) A bounded window over an unbounded stream (Kafka, a log tail, CDC) — peek without hanging
list_sources() What this install can reach, and which are streams

Credentials are host-configured (OMNI_FETCHER_GITHUB_TOKEN, OMNI_FETCHER_POSTGRES_PASSWORD, …) and injected per call — a token never enters the model's context. Point Claude Desktop, Claude Code, or any MCP client at it and the model pulls a Jira ticket, a database row, and a PDF into context through one interface.

Two agent skills ship too: one teaches an agent the contract, one assembles context from a set of sources.

How it compares

Honest positioning — these tools are good at what they do; they solve a differently-shaped problem:

Great at Where OmniFetcher differs
LangChain / LlamaIndex loaders Enormous loader breadth, deep framework integration Each loader returns its own document/metadata shape and raises on failure. OmniFetcher emits one typed contract across every source, returns failures as values, and is framework-free — no chain, no index, no agent runtime to adopt
unstructured.io Best-in-class document parsing into elements Parsing files is one slice. OmniFetcher spans APIs, SaaS, databases, and live streams under the same shape (and will happily hand a PDF to a parser like theirs)
Airbyte / Singer Production ELT — scheduled syncs into a warehouse That's deployed infrastructure and batch pipelines. OmniFetcher is an in-process library for on-demand, request-time fetches: an agent's context, a RAG build, an internal endpoint
Just calling each SDK Total control, zero abstraction Perfect for one or two sources. The cost is N auth models, N response shapes, and N failure modes — to learn, and to maintain

Why it's built this way

The connectors are interchangeable plumbing. The typed shape they emit is the point — and these are the bets behind it, all tested:

  • Errors are values, not exceptions. A missing resource, a bad credential, a parse failure comes back as a typed Error(kind) you branch on — not a try/except you forgot to write. Partial data comes back as Partial with typed gaps naming exactly what's missing, so you never silently lose the half that worked.
  • Stateless, multi-tenant by construction. Connectors hold no credentials and no data between calls; you pass auth per call. One shared orchestrator serves every tenant concurrently — proven by a test that drives it from 16 threads and 64 interleaved coroutines and asserts zero cross-tenant leakage.
  • Read-only means read-only — enforced by the engine, not a regex. The SQL connectors run every query inside a READ ONLY transaction (Postgres/MySQL) or a mode=ro open (SQLite). A write is refused by the database, not by us trying to parse your SQL. Verified against real Postgres, MySQL, and MariaDB.
  • Zoom is semantic, not token-chopping. Ask for text at paragraph or sentence depth and you get the source's own structure, split losslessly — never an arbitrary character window. Ask for a depth a source can't provide and you get an honest gap, not a silent no-op.

The full rationale — including what OmniFetcher deliberately refuses to be — is in PHILOSOPHY.md.

The contract

Every fetch returns a Result. That's the whole surface you branch on:

Result
├── Success ── tree: CompositionNode
├── Partial ── tree + gaps: list[Gap]        # what came back, plus typed holes
└── Error ──── kind: ErrorKind + message     # returned as a value, never raised

CompositionNode          # exactly two fields — no surprises
├── metadata: Metadata   # id, kind, tags, created, updated, author, source_url
│                        #   + source_extra["<source>"] for the source-only fields
└── children: [ ... ]    # a MIXED list, in document order:
                         #   CompositionNode | Text | Image | Audio | Video | Table

Content and metadata are separate on purpose: fields every source has (id, author, timestamps, source_url) are typed on Metadata; anything only one source has is namespaced under source_extra["jira"], so it can't collide with your generic code and can't leak into it.

30 sources — and counting

Documents & files — PDF, DOCX, PPTX, CSV, audio, local files, any URL

Local paths and file://, .pdf, .docx/.pptx (the office extra), .csv, audio, plus any https:// page or JSON/GraphQL/RSS endpoint.

SaaS & collaboration — Jira, Confluence, Slack, Notion, Linear, GitHub, Google Drive, SharePoint, YouTube

Each with its own URI scheme (jira://issue/KEY, slack://channel/ID, …) and per-call auth. The Atlassian ones need the jira / confluence extra.

Databases — Postgres, MySQL/MariaDB, SQLite, Postgres CDC

postgres://, mysql:///mariadb://, sqlite:// run a read query and return a Table atom, read-only enforced by the engine. postgres-cdc:// streams row-level INSERT/UPDATE/DELETE via logical replication. SQLite needs no extra.

Search & storage — Elasticsearch, S3

es://… drives the scroll API and returns one document per hit; s3://bucket/key with AwsAuth.

Streams — Kafka, Redis Streams, log tails, WebSocket, SSE

Unbounded sources consumed with stream() (or the MCP sample tool). Each item carries its resume position, and stream_with_restart reconnects across a dropped connection without loss.

Don't trust a list in a README — ask the installed package what it can actually reach (only sources whose extras are installed show up):

omni-fetcher v1 sources          # a table of every routable source + bounded/stream flag
omni-fetcher v1 sources --json   # the same, as JSON

60 seconds

import asyncio
from omni_fetcher.v1 import AtomKind, Success
from omni_fetcher.v1.connectors.local_file import LocalFileFetcher

async def main() -> None:
    result = await LocalFileFetcher().fetch("README.md")
    if isinstance(result, Success):
        for atom in result.tree.find_atoms(AtomKind.TEXT):
            print(atom.content[:100])

asyncio.run(main())

Optional extras gate a source or capability: office, jira, confluence, elasticsearch, kafka, websockets, postgres, mysql, gdrive, web, mcp. Everything else works on the core install.

A few more moves

Handle all three states — expected failures are values:

result = await connector.fetch(uri)
if isinstance(result, Success):
    ...                       # result.tree
elif isinstance(result, Partial):
    ...                       # result.tree + result.gaps  (don't throw this away)
elif isinstance(result, Error):
    ...                       # result.kind, result.message

Auth per call — nothing is stored, which is what makes it multi-tenant-safe:

await JiraConnector().fetch("jira://issue/PROJ-1",
                            auth=BasicAuth(username="dev@acme.io", password="api-token"))
await SlackConnector().fetch("slack://channel/C0123456789",
                             auth=BearerAuth(token="xoxb-..."))

Consume a stream — resumable across drops:

async for item in omni.stream("tail:///var/log/app.log?from=end"):
    if isinstance(item, Success):
        print(item.tree.find_atoms(AtomKind.TEXT)[0].content)

Or skip Python — the CLI takes credentials as env-var names, so no secret touches argv:

omni-fetcher v1 fetch "sqlite:///app.db?query=SELECT%20*%20FROM%20users" --json
omni-fetcher v1 fetch "jira://issue/PROJ-1" \
  --auth-type basic --username-env JIRA_USER --password-env JIRA_TOKEN
omni-fetcher v1 stream "kafka://localhost:9092/events?offset=earliest" --json --max-items 100

Write a connector — subclass BaseFetcher, override stream(), and fetch() comes free. Any built-in's exact URI/auth/options live in its module docstring (python -c "from omni_fetcher.v1.connectors import mysql; print(mysql.__doc__)"), which never drifts from the code.

Status

v1.10, MIT, ~1,300 tests, Python 3.10–3.12. The v1 contract is stable and additive — new connectors slot in without changing the shape you code against. Issues and connector contributions welcome; if a source you need is missing, the custom-connector path above is about twenty lines.

Development

git clone https://github.com/Jainil-Gosalia/omni_fetcher
cd omni_fetcher
pip install -e ".[dev,office,web,gdrive,confluence,slack,jira,kafka,mcp,mysql]"

pytest --ignore=tests/integration     # unit + contract suites
ruff check . && ruff format --check . && mypy omni_fetcher/

CI runs the same gates on Python 3.10–3.12. Releases are tag-driven: a GitHub Release vX.Y.Z builds, gates, and publishes to PyPI via trusted publishing. The pre-1.0 API still ships unchanged; docs/migration-v1.md maps it onto the contract above.

License

MIT © 2024–2026 OmniFetcher Contributors

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

omni_fetcher-1.11.0.tar.gz (553.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

omni_fetcher-1.11.0-py3-none-any.whl (355.3 kB view details)

Uploaded Python 3

File details

Details for the file omni_fetcher-1.11.0.tar.gz.

File metadata

  • Download URL: omni_fetcher-1.11.0.tar.gz
  • Upload date:
  • Size: 553.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for omni_fetcher-1.11.0.tar.gz
Algorithm Hash digest
SHA256 2fdbfd3128fb1ccf410438b997b6141454710f6818ad7e95b71047c4a716310b
MD5 400f77586d64d826f4f89c8ade5592a0
BLAKE2b-256 a220a03f549a64000ad1121b51b81dca669af11186205b5d2abc557367a81bc7

See more details on using hashes here.

Provenance

The following attestation bundles were made for omni_fetcher-1.11.0.tar.gz:

Publisher: publish.yml on Jainil-Gosalia/omni_fetcher

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file omni_fetcher-1.11.0-py3-none-any.whl.

File metadata

  • Download URL: omni_fetcher-1.11.0-py3-none-any.whl
  • Upload date:
  • Size: 355.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for omni_fetcher-1.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a76322a23cc60c12c25672fa65c09692a438da7c837f32613a7d93b705d5597a
MD5 99f30f37a5329601e45ca1cb47b1af2b
BLAKE2b-256 cde210a2bea442bad5dcf880c6af1d10a0ee2a2202ca81eb6ec6e2fb96fefade

See more details on using hashes here.

Provenance

The following attestation bundles were made for omni_fetcher-1.11.0-py3-none-any.whl:

Publisher: publish.yml on Jainil-Gosalia/omni_fetcher

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page