Skip to main content

kbforge

PyPI CI Python 3.12+ License: MIT

Agent-first knowledge bases, forged from your systems of record.

The Open Knowledge Format (OKF) v0.2 standardizes the artifact at rest — markdown concept files, frontmatter, index.md, log.md. It says nothing about how those bundles get produced: how you pull from a wiki or a CMDB, how you tell a real change from an export timestamp jittering, how a claim stays traceable to its source, and how an update reaches main without a human losing an afternoon to review.

kbforge is the missing half — the production protocol.

Layer Standardized by
Artifact format OKF v0.2
Production protocol — connectors, canonicalization, diff, provenance, publish kbforge
Serving protocol MCP — or any context database that ingests the bundle

"Agent-first" is a checkable claim, not a downstream hope. kbforge stays a producer — the agent connects over MCP, which kbforge doesn't own — but every publish is gated on four agent-facing artifact laws (facet well-formedness, link resolvability, anchor presence, freshness legibility), plus a projection↔files coherence check so nothing ships unvalidated. That gate is what puts the frontmatter, links, and provenance an agent's serving layer needs into the artifact. What each law enforces at full versus reduced strength (and the paths to full strength) is spelled out honestly in architecture.md §4.4 and the artifact-contract spec §5.1.

Status

Alpha — a working walking skeleton. The deterministic core runs end to end with no credentials: two built-in connectors (local_files, git_commits), canonicalization with a stability law, a replay-safe mirror and diff, the §4.4 validator gate, and a dry-run publisher, plus change detection, the no-op rule, and incremental sync via a real cursor, all exercised by the test suite. Two credentialed publishers, GitHub and GitLab, are also available (opt-in via --publisher, token from an env var). Synthesis ships in two forms: a deterministic stub (the default, no LLM) and an opt-in grounded LLM synthesizer (--synthesizer llm, via the kbforge[llm] extra).

Not built yet: a credentialed system-of-record connector. See docs/architecture.md for the full map.

Quickstart

pip install kbforge
kbforge list                       # show available connectors

kbforge run \
  --connector local_files \
  --set path=./docs \
  --mirror .kbforge/mirror --out .kbforge/out --state .kbforge/state

Re-running with no source change is a no-op — no merge request is opened. Point --connector git_commits --set repo=. at a git repository to sync commit history incrementally instead. Config values are YAML-typed, so --set max_commits=50 is an integer and --set 'ignore_globs=[drafts]' is a list.

To synthesize real prose instead of the deterministic stub, install the LLM extra and select the synthesizer (config values are YAML-typed; the API key comes from an env var, never the CLI):

pip install "kbforge[llm]"
export OPENROUTER_API_KEY=...        # or point --llm-set api_base=... at a gateway
kbforge run --connector local_files --set path=./docs \
  --synthesizer llm --llm-set model=deepseek/deepseek-v4-flash \
  --mirror .kbforge/mirror --out .kbforge/out --state .kbforge/state

The synthesizer reaches models through a LiteLLM provider, so OpenRouter and a self-hosted LiteLLM gateway share one config path.

Publishing to GitHub or GitLab

The default publisher writes the proposal to a local directory. To open a real pull request or merge request instead, select a forge publisher and give it a repo. The token comes from an env var, never the CLI:

export GITHUB_TOKEN=...            # or GITLAB_TOKEN
kbforge run --connector local_files --set path=./docs \
  --publisher github --publish-set repo=acme/knowledge-base \
  --mirror .kbforge/mirror --out .kbforge/out --state .kbforge/state

Both publishers accept the same config: repo (required), base (default: the repo's default branch), base_path (a subdirectory, default: repo root), branch (default: sync/<system>), title, api_base (point it at GitHub Enterprise or a self-managed GitLab), and token_env.

kbforge maintains one long-lived sync branch and one open review request per source system: a later run force-updates that branch and edits the existing PR/MR rather than opening a second one. Three consequences worth knowing:

  • Concepts deleted from the source are deleted from the target repo, provided the connector emits an explicit tombstone. Absence never implies a deletion.
  • Manual commits on the sync branch are preserved while its review request is open — a later run builds on the branch rather than resetting it. Once no request is open, the next run rebuilds the branch from the default branch and those commits are gone. A hand edit to a concept kbforge later regenerates is overwritten by that regeneration either way.
  • Close a kbforge review request only by merging it. The mirror advances on every successful publish, so the concepts a request carries are never re-proposed. Closing one unmerged discards its contents permanently: the target repo simply never gets them, and a published-then-abandoned deletion leaves the doc gone from the mirror, so no later run even sees it as a removal. To undo an abandoned request, reset both the mirror and the connector's cursor: delete the mirror directory and, in the state directory (--state), the connector's cursor-<connector-name>.json. Deleting the mirror alone does not work for an incremental connector — its cursor still points past the abandoned content, so the next kbforge_fetch returns only what changed since then, which can be little or nothing, and no re-proposal happens at all. Only once both are gone does a re-run re-propose everything from scratch.

kbforge never merges. No publisher has a merge method.

Design stance

The core ships zero credentialed connectors and zero CI logic. The two built-in connectors need no credentials and serve as references; real systems of record are plugins, discovered through the kbforge.connectors (and kbforge.publishers) entry-point group without editing kbforge — deployments are separate, private repositories. The interface is the product.

# in a third-party package's pyproject.toml — discovered automatically once installed
[project.entry-points."kbforge.connectors"]
myservice = "my_package:connector"

A complete worked example — a credentialed GitHub Issues connector (~160 lines) with token auth, pagination, and a real incremental cursor — is in examples/github-issues-connector/.

The pipeline order — fetch → normalize → mirror → diff → scope → synthesize → validate → publish — is deliberately not pluggable, and neither are the no-op rule or the never-auto-merge rule. Those are the trust guarantees; making them pluggable would make them optional. Plugins extend stages. They cannot reorder or remove them.

Documentation

Related projects

kbforge is one of three contracts for agents, split by seam:

  • ai-agent-contracts — the formal spine: resource, temporal, and lifecycle contracts (the seven-tuple kbforge maps onto).
  • agentic-data-contracts — the consumption half for structured data: domain-driven governance enforced at query time. kbforge is the production half for unstructured knowledge; both independently converged on making freshness legible to the agent.

Not a context database

OpenViking and its kin sit in the serving row of the table above, not the production row. They ingest documents and expose them to an agent — OpenViking summarizes each one into retrieval tiers on write and serves them over a filesystem API and MCP. kbforge produces the documents such a system serves: an OKF bundle in a git repo is a valid input to one, so these compose rather than compete.

The difference shows on the second pull from a source that mostly did not change. A context database refreshes a watched resource by re-ingesting it wholesale — no diff, no changed-set, no proposal a human ever sees. kbforge canonicalizes first, so an export whose timestamps and ordering jitter reduces to no change at all: no LLM spend, no review request, no merge. Byte-level deduplication further downstream cannot substitute, because a rendered system-of-record export is rarely byte-identical across pulls even when nothing about it has meaningfully changed.

That one mechanism is why both the token bill and the review queue stay bounded on a corpus where most documents are stable — and it is what makes the human gate affordable rather than ceremonial. Reach for a context database when an agent needs to retrieve from a corpus; reach for kbforge when a corpus has to stay honest to a system of record that keeps changing, and someone has to be accountable for what it says.

Development

uv sync --all-extras --dev   # create the venv and install
prek install                 # ruff + ty on every commit
uv run pytest

The default suite never touches the network. Tests that call a real external service are marked live and skipped unless you pass --run-live.

The forge publishers have a live suite because their offline tests can only assert what we meant to send — a real forge is the only thing that can say the intent was right. It needs a throwaway repo on each forge and the two CLIs (gh, glab) authenticated; each run writes under a fresh live/<run-id>/ prefix, so nothing accumulates and no repo is ever deleted.

GITHUB_TOKEN=$(gh auth token) \
GITLAB_TOKEN=$(glab config get token --host gitlab.com) \
KBFORGE_LIVE_GITHUB_REPO=you/kbforge-live-test \
KBFORGE_LIVE_GITLAB_REPO=you/kbforge-live-test \
uv run pytest tests/test_forge_live.py --run-live

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kbforge-0.5.0.tar.gz (199.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kbforge-0.5.0-py3-none-any.whl (51.7 kB view details)

Uploaded Python 3

File details

Details for the file kbforge-0.5.0.tar.gz.

File metadata

  • Download URL: kbforge-0.5.0.tar.gz
  • Upload date:
  • Size: 199.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for kbforge-0.5.0.tar.gz
Algorithm Hash digest
SHA256 11de31299d4de63f7df931f98bedf1c12776d648436ba9ed5bd8145af227b0af
MD5 bb258a1f52cec12e4739a0823c3378af
BLAKE2b-256 96e0a63fb7cdaa3d64f7c232939e4fc06a22ac9b0b96cf176a3719162db33793

See more details on using hashes here.

Provenance

The following attestation bundles were made for kbforge-0.5.0.tar.gz:

Publisher: ci.yml on flyersworder/kbforge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file kbforge-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: kbforge-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 51.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for kbforge-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0bc8d0a8f4c7b872ff99ca969acd9a149327c796bf15c31ee7d443abf425b9e3
MD5 f1abc6ba9b6039fd53176f4babf5b4ad
BLAKE2b-256 0436d8b647cd48e91c1d83879b29130cd5495c4243c5bf5125a00fbd20c95863

See more details on using hashes here.

Provenance

The following attestation bundles were made for kbforge-0.5.0-py3-none-any.whl:

Publisher: ci.yml on flyersworder/kbforge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page