Skip to main content

DocForge SDK — INGESTION, FORGED

The typed Python client for DocForge — async and sync, fully type-hinted, zero server-tree dependency.

PyPI Python License


docforge-sdk is a typed Python client for the DocForge REST API — a document-intelligence platform that melts any document (PDF, Office, images…) into a canonical intermediate representation, enriches and chunks it, embeds it, and serves hybrid retrieval over it.

The SDK ships both an asynchronous and a synchronous client with an identical surface, is fully type-hinted (py.typed, Pydantic v2 models), and has zero dependency on the DocForge server tree — it is a clean-room client that talks to the API over HTTP only (httpx + pydantic + hand-written models mirroring the public REST contract), so it can be vendored or published independently.

  • Async + sync — the same methods with and without await.
  • Typed end to end — request/response Pydantic models, no dict-spelunking.
  • Context-managed — connection pooling via async with / with.
  • Typed errors — one exception hierarchy for connection, auth, 404, 409, 422…
  • Tiny footprint — only httpx and pydantic.

Table of contents


Install

pip install docforge-sdk

Requirements

  • Python 3.12+
  • A running DocForge API (self-hosted). The examples assume http://localhost:10040.
  • An API token when the server has auth enabled (see Authentication).

Quickstart

Async

import asyncio

from docforge_sdk import AsyncClient, SearchRequest


async def main() -> None:
    async with AsyncClient("http://localhost:10040", api_token="df_root_...") as client:
        # Liveness
        print(await client.health.ping())

        # List collections
        collections = await client.collections.list()
        for c in collections:
            print(c.id, c.name)

        # Search the first one
        hits = await client.search.search(
            collections[0].id,
            SearchRequest(query="quarterly revenue", limit=5),
        )
        for hit in hits.hits:
            print(f"{hit.score:.3f}  {hit.text[:80]}")


asyncio.run(main())

Sync

The synchronous client is the async one without await — same resources, same signatures.

from docforge_sdk import Client, SearchRequest

with Client("http://localhost:10040", api_token="df_root_...") as client:
    collections = client.collections.list()
    hits = client.search.search(
        collections[0].id,
        SearchRequest(query="quarterly revenue", limit=5),
    )
    for hit in hits.hits:
        print(hit.score, hit.text)

Both clients accept the same constructor:

AsyncClient(base_url: str, timeout: float = 30.0, api_token: str = "")
Client(base_url: str, timeout: float = 30.0, api_token: str = "")
Argument Default Meaning
base_url API origin, e.g. "http://localhost:10040".
timeout 30.0 Per-request timeout in seconds.
api_token "" Bearer token; empty means unauthenticated requests.

Always use the client as a context manager (async with / with) so the underlying HTTP connection pool is opened and closed cleanly. You can construct it directly, but then you own await client.aclose() / client.close().

Authentication

When the DocForge server runs with auth enabled, pass a bearer token as api_token. The token is sent as Authorization: Bearer <token> on every request.

client = AsyncClient("https://docforge.example.com", api_token="df_...")

Keys are scoped: a key carries a set of capabilities (read, write, search, create, admin) and an optional allow-list of collection ids. The root token created at server bootstrap has full access; mint narrower keys with client.auth (see Manage API keys).

A key holding create may create new collections, and is auto-granted ownership of what it creates: the new collection's id is appended to that key's own scope, so a single key can be given "may create collections + full power over the ones it creates" without knowing ids in advance — the natural setup for driving DocForge from an agent (e.g. over MCP). Pair it with a narrow, search-only key scoped to one collection for your app's runtime.

Resources & methods

Every resource hangs off the client (client.<resource>.<method>(...)). The async and sync surfaces are identical.

health

Method Returns Purpose
ping() HealthStatus Liveness probe.

collections

Method Returns Purpose
list() list[CollectionListItem] Every collection (schema + pipelines), each with its server-computed health summary.
get(collection_id) CollectionModel One collection (schema + pipelines).
create(CreateCollectionRequest) CollectionModel Create a collection (contract).
update(collection_id, UpdateCollectionRequest) CollectionModel Patch name / formats / fields / pipelines / cost-estimate overrides.
delete(collection_id) None Delete a collection.
health(collection_id) CollectionHealthResponse On-demand operational health — 5-state verdict, provider reachability sweep, index stats. No job enqueued, no spend.
storage(collection_id) CollectionStorageResponse Material storage footprint per store (S3 exact, Postgres/Qdrant estimated) + per-document breakdown.
estimate(collection_id, scope="pending", document_ids=None, filter=None) CostEstimate Pre-hoc cost + volume dry-run over the whole collection (scope), a selected subset (document_ids), or a corpus filter. No spend. Per-collection rate/assumption overrides apply.
contract_schema() CollectionContractSchemaResponse JSON Schema of the identity/limits contract (build a valid create/update).

documents

Method Returns Purpose
upload(collection_id, file, metadata=None, filename=None) UploadAccepted Admit a document; returns the ingestion job_id. file is a path, bytes, or a Path.
set_enabled(document_id, enabled) DocumentEnabledResponse Reversibly hide/show a document from search.
get_markdown(document_id, download=False) DocumentView The document rendered as Markdown from the canonical IR (a generated view).
get_html(document_id, download=False) DocumentView The document rendered as HTML from the canonical IR (a generated view).
reingest(document_id, force=False) UploadAccepted Re-run the full ingestion of a single document.

search

Method Returns Purpose
search(collection_id, SearchRequest) SearchResponse Hybrid (dense + sparse RRF) search, optional rerank.

explorer

Method Returns Purpose
list_documents(collection_id) list[DocumentListItem] Documents in a collection.
get_document(document_id) DocumentDetail One document's detail.
get_pages(document_id) list[PageInfo] Page-level info.
get_ir(document_id) DocumentIRModel The canonical IR (blocks, tables, figures, enrichments).
get_chunks(document_id) list[ChunkInfo] The document's chunks.
delete_document(document_id) None Delete a document.
set_chunk_enabled(chunk_id, enabled) ChunkEnabledResult Toggle one chunk in/out of search.
set_chunks_enabled(BulkChunkEnabledPatch) BulkChunkEnabledResponse Toggle many chunks at once.

jobs

Method Returns Purpose
list(collection_id) list[JobStatus] Ingestion jobs for a collection.
get(job_id) JobStatus One job's status + progress.
get_events(job_id) JobTrace Per-stage event trace.
live_workers() WorkersLive Currently active workers.
cancel(job_id, force=False) CancelResult Stop a job — cooperative by default, immediate with force=True.
cost(collection_id) CollectionCost Paid text-gen roll-up (tokens + USD).
queue(collection_id=None) QueueDepth Pending/running backlog (fleet-wide or per-collection).
stage_durations(collection_id) StageDurations Average per-stage wall-clock (ETA basis).

blobs

Method Returns Purpose
get(content_hash) BlobContent Fetch a content-addressed blob (bytes + media type).

corpus

Server-side document grid + bulk operations at scale.

Method Returns Purpose
query(collection_id, DocumentQueryRequest) DocumentQueryResponse Filtered/sorted/paginated page + total match count.
bulk_delete(collection_id, DocumentSelector) BulkDeleteResponse Delete by id-set XOR filter-minus-excludes.
bulk_set_enabled(collection_id, DocumentSelector, enabled) BulkEnabledResponse Bulk enable/disable searchability.
bulk_reingest(collection_id, DocumentSelector, force=False) BulkReingestResponse Bulk full re-ingest (capped fan-out).

snippets

Granular, config-only export/import of ONE collection slice — synchronous, secret-masked, versioned (saved as *.dfsnippet, distinct from the whole-collection .dcexport bundle).

Method Returns Purpose
export(collection_id, kind) CollectionSnippet Export the ingestion pipeline, the search graph, or the metadata schema as a portable snippet.
apply(collection_id, kind, snippet) SnippetImportResult Apply a snippet of that kind onto an existing collection (healed/validated like a PATCH).

auth

Method Returns Purpose
create_key(name, permissions=None, expires_at=None) CreatedKey Mint a key. The plaintext token is returned once.
list_keys() list[KeyInfo] Every key (metadata only, never the secret).
rotate_key(key_id, ...) CreatedKey Roll a key's secret.
revoke_key(key_id) None Revoke a key.
whoami() WhoAmI THIS token's own capabilities + collection scope.

pipelines

Discovery + design surface for the ingestion / search graphs (advanced).

Method Returns Purpose
list_surfaces() PipelineIndexResponse Available pipeline designs.
get_design(key, full=True) PipelineDesignResponse A design's palette + default blob.
inspect(key, blob) InspectResponse Validate a graph blob.
edit(key, ...) EditResponse Apply a structural edit to a blob.
view_stages(key, blob) StageViewResponse Compile a blob into the stage-rail view.
apply_stage(key, ...) StageApplyResponse Apply one stage-rail action.

Recipes

Create a collection

A collection is a contract: a metadata schema + an ingestion pipeline + a search pipeline. The pipelines default to the stock graphs when omitted.

from docforge_sdk import Client, CreateCollectionRequest, FieldSpec, FieldType

with Client("http://localhost:10040", api_token="df_...") as client:
    collection = client.collections.create(
        CreateCollectionRequest(
            name="reports",
            supported_formats=["pdf", "docx"],
            max_file_size_bytes=50 * 1024 * 1024,
            fields=[
                FieldSpec(field_name="year", field_type=FieldType.INTEGER, filterable=True),
                FieldSpec(field_name="team", field_type=FieldType.KEYWORD_LIST, filterable=True),
            ],
        )
    )
    print(collection.id)

Key CreateCollectionRequest fields beyond the schema: max_file_size_bytes (bytes) and job_timeout_seconds (float | None, seconds) — the whole-ingest-job wall-clock budget for that collection; None (the default) inherits the worker's global job-timeout default. Same field, same semantics on CollectionModel (read) and UpdateCollectionRequest (write; there, omitting it leaves the current value unchanged, a set value overrides it).

Upload a document and wait for ingestion

upload returns immediately with a job_id; ingestion runs asynchronously on the worker. Poll the job until it reaches a terminal state.

import time
from docforge_sdk import Client

with Client("http://localhost:10040", api_token="df_...") as client:
    accepted = client.documents.upload(
        collection_id,
        file="report-2024.pdf",
        metadata={"year": 2024, "team": ["finance"]},
    )

    while True:
        job = client.jobs.get(accepted.job_id)
        print(job.status, job.progress, job.current_stage)
        if job.status in {"done", "failed"}:
            break
        time.sleep(2)

    if job.status == "failed":
        raise RuntimeError(job.error)

The async variant is the same with await and asyncio.sleep.

Hybrid search with filters

SearchRequest drives dense + sparse fusion (RRF), optional late-interaction rerank, and metadata filtering.

from docforge_sdk import Client, SearchRequest

with Client("http://localhost:10040", api_token="df_...") as client:
    resp = client.search.search(
        collection_id,
        SearchRequest(
            query="revenue guidance for next fiscal year",
            limit=10,
            filters={"year": 2024, "team": "finance"},
        ),
    )
    for hit in resp.hits:
        print(f"{hit.score:.3f}  doc={hit.document_id}  {hit.text[:100]}")

Key SearchRequest fields: query (required), limit (1–100, default 10), filters (dict[str, Any]), search_in (list[SearchTarget] to pick which vectors to query).

Explore a document (pages, IR, chunks)

with Client("http://localhost:10040", api_token="df_...") as client:
    docs = client.explorer.list_documents(collection_id)
    doc = client.explorer.get_document(docs[0].document_id)

    ir = client.explorer.get_ir(doc.document_id)  # canonical IR
    chunks = client.explorer.get_chunks(doc.document_id)

    # Reversibly drop a noisy chunk out of search:
    client.explorer.set_chunk_enabled(chunks[0].chunk_id, enabled=False)

Measure a collection's storage footprint

storage reports the material footprint per store — S3 bytes are exact (deduped), Postgres/Qdrant bytes are estimates (each section flags this via its own estimated) — plus a per-document breakdown sorted heaviest first.

with Client("http://localhost:10040", api_token="df_...") as client:
    footprint = client.collections.storage(collection_id)
    print(footprint.grand_total_bytes, footprint.s3.physical_unique_bytes)
    for doc in footprint.documents[:5]:
        print(doc.filename, doc.total_bytes)

Manage API keys

Mint a scoped, expiring key (the plaintext is shown once, on creation):

from datetime import datetime, timedelta, timezone
from docforge_sdk import Client, Capability, KeyPermissions

with Client("http://localhost:10040", api_token="df_root_...") as client:
    created = client.auth.create_key(
        name="reporting-bot",
        permissions=KeyPermissions(
            capabilities=[Capability.READ, Capability.SEARCH],
            collections=[collection_id],  # empty list = all collections
        ),
        expires_at=datetime.now(timezone.utc) + timedelta(days=90),
    )
    print("SAVE THIS NOW:", created.key)  # plaintext, only returned once

    for key in client.auth.list_keys():
        print(key.id, key.name, key.last_used_at)

Error handling

Every failure raises a subclass of DocForgeError, so you can catch broadly or precisely.

from docforge_sdk import (
    Client,
    DocForgeError,
    APIConnectionError,
    APITimeoutError,
    APIStatusError,
    AuthError,
    NotFoundError,
    ConflictError,
    UnprocessableError,
)

with Client("http://localhost:10040", api_token="df_...") as client:
    try:
        client.collections.get("does-not-exist")
    except NotFoundError:
        ...  # 404
    except AuthError:
        ...  # 401 / 403 — bad or unscoped token
    except UnprocessableError as e:
        ...  # 422 — validation errors (see e.status / e.detail)
    except APITimeoutError:
        ...  # request exceeded `timeout`
    except APIConnectionError:
        ...  # server unreachable
    except DocForgeError:
        ...  # catch-all

Hierarchy:

DocForgeError
├── APIConnectionError
│   └── APITimeoutError
└── APIStatusError            # any non-2xx (carries .status)
    ├── AuthError             # 401 / 403
    ├── NotFoundError         # 404
    ├── ConflictError         # 409
    └── UnprocessableError    # 422

Type hints & discoverability

The package ships py.typed, so editors and type-checkers see every request/response type. All public models are re-exported from the top level:

from docforge_sdk import (
    # clients
    AsyncClient,
    Client,
    # collections
    CollectionModel,
    CreateCollectionRequest,
    UpdateCollectionRequest,
    FieldSpec,
    FieldType,
    CollectionStorageResponse,
    # documents / explorer
    UploadAccepted,
    DocumentDetail,
    DocumentListItem,
    ChunkInfo,
    PageInfo,
    DocumentIRModel,
    # search
    SearchRequest,
    SearchResponse,
    SearchHit,
    SearchTarget,
    # jobs
    JobStatus,
    JobTrace,
    JobEvent,
    WorkersLive,
    # auth
    Capability,
    KeyPermissions,
    CreateKeyRequest,
    CreatedKey,
    KeyInfo,
    # errors
    DocForgeError,
    AuthError,
    NotFoundError,
    ConflictError,
    UnprocessableError,
)

from docforge_sdk import * also works and pulls the full public surface.

Versioning & compatibility

The SDK tracks the DocForge REST contract; a CI parity gate diffs the SDK models against the live server's OpenAPI on every change, so a published version is coherent with the API it targets. Pin a version in production:

pip install "docforge-sdk==0.3.0"

License

MIT — see LICENSE. This SDK is deliberately licensed MIT even though the parent DocForge repository is GPLv3: it is a standalone, clean-room client (HTTP models only, no server code), so a permissive per-directory license is intentional and lets any project depend on it freely.

Publishing (maintainers)

Releases publish to PyPI via Trusted Publishing (OIDC) — there is no API token stored in the repo. To cut a release, tag a commit with the sdk-v<version> prefix (the version must match docforge_sdk/_version.py) and push the tag:

git tag sdk-v0.1.0
git push origin sdk-v0.1.0

The .github/workflows/release-sdk.yml workflow then builds and uploads the sdist + wheel.

One-time PyPI setup (done once by the maintainer, before the first release):

  1. Reserve the project name docforge-sdk on PyPI.
  2. Under the project's Publishing settings, add a GitHub trusted publisher with:
    • Owner: Florian-BARRE
    • Repository: docforge
    • Workflow name: release-sdk.yml
    • Environment: pypi

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docforge_sdk-0.14.6.tar.gz (165.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

docforge_sdk-0.14.6-py3-none-any.whl (88.5 kB view details)

Uploaded Python 3

File details

Details for the file docforge_sdk-0.14.6.tar.gz.

File metadata

  • Download URL: docforge_sdk-0.14.6.tar.gz
  • Upload date:
  • Size: 165.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docforge_sdk-0.14.6.tar.gz
Algorithm Hash digest
SHA256 558b0ee3f43cbc804b6de3a15c858a2a70e872b7209e2ba0b444e574a4d3ac27
MD5 5c493efb4cf8f503969ed06e5c30daf1
BLAKE2b-256 4ed5cceb6c318a45aa4dca7fa22ca50594ea356bd44382fb3237a01331349af8

See more details on using hashes here.

Provenance

The following attestation bundles were made for docforge_sdk-0.14.6.tar.gz:

Publisher: release-sdk.yml on Florian-BARRE/DocForge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file docforge_sdk-0.14.6-py3-none-any.whl.

File metadata

  • Download URL: docforge_sdk-0.14.6-py3-none-any.whl
  • Upload date:
  • Size: 88.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docforge_sdk-0.14.6-py3-none-any.whl
Algorithm Hash digest
SHA256 ea82e68530b509cdc6ed2860b7438194f37174883fe0af9e78efd2881f4d9209
MD5 bc7473e9323a864c164f270c12fc2c5b
BLAKE2b-256 394537941deadbc3ef51d37b0b009e3173ad75738af9a110ebe57185dd847ce2

See more details on using hashes here.

Provenance

The following attestation bundles were made for docforge_sdk-0.14.6-py3-none-any.whl:

Publisher: release-sdk.yml on Florian-BARRE/DocForge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.14.10

2 files

0.14.9

2 files

0.14.8

2 files

0.14.7

2 files

This release

0.14.6 This release

2 files

0.14.4

2 files

0.14.3

2 files

0.14.0

2 files

0.13.0

2 files

0.12.1

2 files

0.9.12

2 files

0.9.11

2 files

0.9.9

2 files

0.9.0

2 files

0.8.1

2 files

0.7.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page