Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

harborrag-adapters

Provider adapters for HarborRAG.

This package owns the code that talks to external systems and extracts text from raw source files. It sits between:

  • harborrag-core, which owns shared domain schemas such as SourceRecord, RawDocument, ParseInput, and ParsedDocument.
  • harborrag-runtime, which should own orchestration, concurrency, scheduling, backpressure, retries across tasks, and large ingestion pipelines.

The adapter layer should stay focused on source-specific behavior: auth, API pagination, request retry hints, payload safety limits, parser routing, and clear errors for missing optional dependencies.

Public extension API

Each adapter family owns its registry and factory at the variation point: connectors, parsers, model providers, and repository backends are registered independently. There is intentionally no cross-family registry because adding one provider must not require editing unrelated adapter families.

File ownership and test doubles

  • connectors/base.py and parsers/common/base.py define contracts. Production implementations belong in provider or format modules under those packages.
  • Connector-specific canonical transformation lives in the provider's document_transform.py; supporting representation logic may be split into a provider-local package such as connectors/confluence/normalization/. It does not belong in runtime or the chunk-refinement provider package.
  • Implemented connectors and parsers do not keep production mock.py modules; their fakes and test doubles belong under tests/.
  • Repository families expose provider-neutral Harbor* contracts plus real provider packages. Their deterministic fakes belong under tests/; production repository packages do not ship mock backends.
  • Provider construction belongs to the registry or factory in that capability; domain schemas remain in harborrag-core, and workflow orchestration remains outside adapters.

Contributor and teammate deliverables

A new adapter contribution should include its implementation and validated config, public exports or registry registration, focused unit/failure/security tests, optional-dependency declarations, and documentation or example env keys. Never commit credentials or use real provider payloads as fixtures.

Install

Base connector support only requires requests plus harborrag-core:

pip install -e packages/harborrag-core -e packages/harborrag-adapters

Install common parser dependencies when parsing office files, images, HTML, and PDFs through the default parser stack:

pip install -e "packages/harborrag-adapters[parsers]"

Install the optional third-party chunking provider used for oversized text refinement:

pip install -e "packages/harborrag-adapters[chunking]"

The implementation is chunking/recursive.py. Inject RecursiveTextRefiner into the engine chunking service. It returns HarborRAG-owned TextSplit values, so engine policy never depends on framework-owned types. Raw Markdown, HTML, JSON, PDF, and Office structure is owned by parsers and canonical normalizers, not by the chunk refinement adapter.

Install advanced PDF backends separately when needed. This extra also includes RapidOCR with its default ONNX Runtime CPU engine:

pip install -e "packages/harborrag-adapters[pdf]"

For a Docling/RapidOCR-only deployment, use the narrower extra:

pip install -e "packages/harborrag-adapters[pdf-docling]"

SQLite database, state, filesystem, and memory repositories are included in the base install. Install only the extras needed by deployed services - control-plane adds the Alembic-managed control plane and tables adds the pyarrow-backed canonical table artifacts:

pip install -e "packages/harborrag-adapters[redis,qdrant,falkordb,postgres,s3,control-plane,tables]"

Main Modules

Module Purpose
harborrag_adapters.connectors Source connectors that discover/load source records and own provider-specific canonical normalization.
harborrag_adapters.parsers Parser factory and format parsers that produce ParsedDocuments.
harborrag_adapters.models Chat, embedding, and reranking clients behind provider-neutral contracts.
harborrag_adapters.repositories Tenant-isolated vector, graph, cache, object, database, and workflow-state repositories.
harborrag_adapters.chunking Chunking strategies used by the ingestion engine.

See the module READMEs for deeper notes:

  • src/harborrag_adapters/connectors/README.md
  • src/harborrag_adapters/connectors/confluence/normalization/README.md
  • src/harborrag_adapters/parsers/README.md
  • src/harborrag_adapters/parsers/pdf/README.md
  • src/harborrag_adapters/models/README.md

Quick Start

Load local files and parse them through the default parser registry:

Run this from the repository root - source_uri is resolved relative to the process working directory.

import asyncio
from pathlib import Path

from harborrag_adapters.connectors import (
    ConnectorQuery,
    LocalFileConfig,
    LocalFileConnector,
)
from harborrag_adapters.parsers import HarborParserFactory, ParseRequest


async def main() -> None:
    connector = LocalFileConnector(
        LocalFileConfig(source_path="docs", allowed_extensions={".md", ".txt"})
    )
    registry = HarborParserFactory().create_registry()

    for record in connector.discover(ConnectorQuery(recursive=True)):
        result = await registry.parse_request(
            ParseRequest(
                source_uri=f"docs/{record.locator}",
                filename=Path(record.locator).name,
                mime_type=record.source_type,
            )
        )
        print(record.locator, result.parser_name, result.engine_name, len(result.text))


asyncio.run(main())

Discovery is synchronous and yields SourceRecord objects (id, source_type, locator, metadata, updated_at, checksum). Parsing is asynchronous: the registry selects a parser family and the family routes to a concrete engine, so parse_request returns both names alongside the extracted text.

Create a connector through the provider registry:

from harborrag_adapters.connectors import HarborConnector, GitHubRepositoryConfig

connector = HarborConnector(
    "github",
    config=GitHubRepositoryConfig(
        owner="example",
        repo="knowledge-base",
        branch="main",
        root_path="docs",
    ),
)

records = list(connector.discover())

Connectors

Available connector providers:

Provider Config class Source
local LocalFileConfig Local file or directory trees.
github GitHubRepositoryConfig GitHub repository files through the REST API.
confluence ConfluenceSpaceConfig Confluence Cloud or Data Center pages, comments, and attachments.
jira JiraProjectConfig JIRA Cloud or Data Center issues, comments, changelog, and attachments.
sharepoint SharePointSiteConfig SharePoint document libraries through Microsoft Graph.

Connector flow:

  1. discover(query) returns lightweight SourceRecord objects.
  2. load(record) fetches one RawDocument.
  3. load_raw_documents(query) combines both steps for simple callers.

The connectors include source-level safeguards: same-origin download checks, provider-aware retry delays, file-size gates, nested collection caps, scoped local path handling, and structured logging.

Parsers

HarborParserFactory().create_registry() builds the registry. It routes by filename suffix and MIME content type. Generic transport MIME types defer to a specific suffix; other conflicting routes raise instead of choosing a parser silently.

Routing happens in two levels. The registry picks a family; the family then owns engine selection, quality checks, fallback, and output normalization. The eight registered families and their extensions:

Family Extensions
text plain text plus common source and config suffixes - .txt, .py, .ts, .yaml, .toml, .sql, .rst, and ~30 more
markup .md, .markdown, .mdx, .html, .htm, .xhtml
structured .json, .jsonl, .ndjson
spreadsheet .csv, .tsv, .xls, .xlsx, .xlsm, .xltx, .xltm
document .docx, .odt, .epub
presentation .pptx, .pptm
image .png, .jpg, .jpeg, .gif, .bmp, .tif, .tiff, .webp - via OCR
pdf .pdf, with PyMuPDF, Docling, LiteParse, MinerU, and PaddleOCR engines

List them at runtime with registry.families(). The PptxParser/DocxParser/ PdfParser-style names are migration aliases in parsers/compat.py, kept out of the package __all__; write new code against the registry and family names above.

Repositories

Repository backends are asynchronous context managers and take a StorageOperationContext on every data operation. The context supplies the tenant boundary; callers must not encode tenant IDs into user keys themselves.

Family Providers Main products
Vector Qdrant Collections, point CRUD, scan, dense and hybrid search.
Graph FalkorDB Node/edge CRUD and bounded subgraph expansion.
Cache Memory, Redis JSON values, tags, counters, compare-and-set, fenced locks.
Object store Memory, filesystem, S3 Streaming bodies, metadata, list/delete, presigned reads.
Database SQLite, PostgreSQL Document/chunk unit of work and transactional outbox.
Workflow state SQLite, Redis Versioned state, checkpoints, leases, fencing tokens.

Use the family clients (HarborVectorDBClient, HarborGraphDBClient, HarborCacheDBClient, HarborObjectStoreDBClient, HarborDatabaseClient, and HarborStateDBClient) for provider-name construction, or instantiate a provider backend directly when configuration is already typed.

Live, non-pytest smoke checks for Redis, FalkorDB, PostgreSQL, Qdrant, and SQLite are documented in tests/repositories/smoke/README.md.

Reliability Boundaries

Adapters should:

  • Validate config and source scope early.
  • Keep provider SDK imports inside provider modules.
  • Respect provider retry and rate-limit hints.
  • Fail clearly when optional parser dependencies are missing.
  • Enforce source-level size and collection limits before loading content into memory.

Runtime should:

  • Own ingestion concurrency and queueing.
  • Own cross-source scheduling and global retry policy.
  • Own chunk fan-out, backpressure, and resumable ingestion jobs.
  • Introduce any future stream or spooled-file contract needed for truly huge documents.

Tests

Package tests live in:

packages/harborrag-adapters/tests/

Run from the repository root:

pytest packages/harborrag-adapters/tests

Run the hermetic repository suite and enforce its coverage target with:

pytest -n 0 packages/harborrag-adapters/tests/repositories/unit \
  --cov=harborrag_adapters.repositories --cov-fail-under=90

Keep connector tests close to provider behavior and parser tests focused on format routing, dependency errors, and parsed output shape.

Release files for harborrag-adapters 2.0.0a1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for harborrag-adapters 2.0.0a1
File Size Uploaded
harborrag_adapters-2.0.0a1.tar.gz 740.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for harborrag-adapters 2.0.0a1
File Interpreter ABI Platform
harborrag_adapters-2.0.0a1-py3-none-any.whl Python 3 none any Details

Total release size: 1.4 MB

Release files / harborrag_adapters-2.0.0a1.tar.gz

Download URL harborrag_adapters-2.0.0a1.tar.gz
Size 740.3 kB
Tags Source
SHA-256 checksum
How to use checksums
bd99c7907167838af01f01813a93e37da628d0f1e667e92a406cd8b734549de7
BLAKE2b-256 checksum
How to use checksums
4e21efdf77ac8de6c90ff8809caef3a71eb2d1ba6c07e88bd211bbd619a7c4dd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release files / harborrag_adapters-2.0.0a1-py3-none-any.whl

Download URL harborrag_adapters-2.0.0a1-py3-none-any.whl
Size 693.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
aeb884fa6fdd3692f9fad3d0d488f10defa5dfc36bb5ce04a1b74b2b0340c25e
BLAKE2b-256 checksum
How to use checksums
96ef0d160441fdf59db73e9db68fde7353248cc7ded6b6bbb7acb1ebd27b717e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2.0.0a1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page