Skip to main content

Rust-based epistemic graph engine for agent-utilities

Project description

epistemic-graph

One durable, Rust-native engine that is a drop-in substrate for graph · vector · SQL · SPARQL/RDF/OWL · time-series · blob · key-value
Every modality is a first-class citizen of one RowSet planner — from a Raspberry Pi to a replicated Raft cluster, from one core.

Version Language License

Honesty first. This README claims only what the code actually does today. Every row in the capability matrix is tagged ✅ supported · 🔶 in-progress · 🗺 roadmap, and the parity roadmap names exactly which gaps are being closed and in which order. If a doc and the code disagree, the code wins — file an issue.

Documentation — the full architecture, the tier/binary map, deployment recipes, the per-interface guides, and the concept registry live in the official documentation.

This is the compute & storage engine for agent-utilities — a standalone Rust service reached out-of-process over MessagePack/UDS (no PyO3), embedded in-process, or spoken to over the Postgres wire protocol. Contributing? See CONTRIBUTING.md.


The thesis: one engine, every modality

A modern agent platform normally needs a graph database and a vector index and a SQL warehouse and a triple-store + reasoner and a time-series DB and a blob store and a full-text index — seven systems, seven copies of the data, six sync pipelines, and a brittle application layer that stitches results back together.

epistemic-graph collapses that rack into one durable engine with one unified query planner. Every modality is a view over the same RowSet algebra, so a single plan can seed candidates from an OWL inference or a SPARQL pattern, filter them with SQL, traverse the graph, re-rank by vector similarity and BM25 text, fuse the two, and run a sandboxed WASM UDF — without ever leaving the engine or marshalling rows back to Python.

It is durable by default: built with the redb feature (folded into every deployment tier), the persist dir is the authoritative source of truth and an acked write survives kill -9 (commit-before-ack). It scales by configuration alone — the same binary family runs as an embedded in-process library on a Pi, a single durable server, or a multi-node Raft cluster with cross-shard transactions.

Drop-in positioning — and the honest parity status

You run today epistemic-graph as a drop-in Current parity
Postgres (psql / BI / ORM) pgwire server: SCRAM/trust auth, simple + extended protocol, pg_catalog/information_schema ✅ read SQL + ✅ graph-table DML · 🗺 arbitrary tables + DDL
Stardog / GraphDB (RDF triple-store + reasoner) RDF dataset over the property graph, SPARQL SELECT, OWL 2 EL⁺/RL reasoning ✅ SELECT + ✅ reasoning · 🔶 ASK/CONSTRUCT/DESCRIBE/UPDATE + /sparql endpoint
Neo4j (property graph) native petgraph core + Cypher MATCH…RETURN + graph algorithms ✅ read traversal & algorithms · 🗺 Cypher writes
Pinecone / Milvus (vector DB) native IVF-PQ + OPQ + SQ8 ANN, persistent, warm-on-start
InfluxDB / TimescaleDB (time-series) native redb TSDB: ASOF, gap-fill, time_bucket, OHLC, decay ✅ primitives · 🔶 time-ops as planner ops
S3 / MinIO (blob) content-addressed streaming CAS, redb-native or S3-backed
SQLite / RocksDB (embedded KV) EmbeddedEngine in-process handle over the same redb rows ✅ embedded graph API · 🗺 generic KV / SQLite-wire surface

The point is convergence, not a checkbox: the modalities share one snapshot, one ACID transaction, one security model, and one planner. See the full capability matrix below for the operation-by-operation truth.


Capability matrix

Legend: ✅ supported (implemented & tested) · 🔶 in-progress (partial or being added now) · 🗺 roadmap (designed, not built). Feature flags are the Cargo features that gate each surface; the tier table shows which prebuilt binary carries them.

Interface Operation Status Feature Notes
SQL SELECT (joins, aggregates, CTE, window, subquery) query DataFusion 43 over nodes + edges; real predicate pushdown (Inexact)
SQL INSERT / UPDATE / DELETE query nodes table only, literal VALUES, single col = literal WHERE (KG-2.198)
SQL Complex/compound WHERE, INSERT…SELECT, JOIN-in-DML 🔶 query KG-2.198 follow-up; explicitly errors today
SQL Arbitrary user tables + DDL (CREATE/ALTER/DROP) 🗺 query errors unsupported statement; being added now
Postgres wire listener, simple + extended/prepared protocol pgwire EPISTEMIC_GRAPH_PGWIRE_ADDR; also pulled in by cluster
Postgres wire SCRAM-SHA-256 / trust auth, pg_catalog introspection pgwire KG-2.202 / KG-2.201; pg user → engine ACL actor
SPARQL SELECT (BGP, paths, FILTER subset, OPTIONAL, UNION, GROUP/agg, BIND, DISTINCT, SLICE) sparql spargebra parser compiled to LPG scans
SPARQL ASK / CONSTRUCT / DESCRIBE 🔶 sparql errors supports SELECT only; being added now
SPARQL UPDATE (INSERT/DELETE DATA) 🗺 sparql absent in eg-rdf
SPARQL /sparql HTTP endpoint 🗺 today exposed as binary RPC Method::Sparql
SPARQL true named graphs, SPO/POS index, regex/arith FILTER 🔶 sparql single default graph; naive full-scan; FILTER is a subset
Cypher MATCH … WHERE … RETURN … LIMIT (var-length single hop) cypher read-only over one snapshot; WHERE is AND-only
Cypher writes (CREATE/MERGE/SET/DELETE), ORDER BY/WITH/OR 🗺 cypher not in grammar
GraphQL read queries (scan + BFS, schema-from-graph, aliases, first/limit, filters) graphql byte-equal to Cypher path
GraphQL mutations / subscriptions / fragments / variables 🗺 graphql explicitly rejected at parse
OWL EL⁺ + RL forward-chaining materialization & classification owl pure-Rust; consistency + incremental + justifications
OWL confidence-weighting + Ebbinghaus time-decay owl KG-2.236; per-axiom eg:confidence, fact decay
OWL query-time Op::Reason (reasoner seeds a RowSet) owl-plan distributed/cross-shard union supported
OWL OWL-DL (tableau, cardinality, allValuesFrom), SWRL user rules 🗺 out of the EL+RL envelope by design
Vector / ANN IVF-PQ + OPQ + SQ8-refine, persistent (reopen w/o rebuild), warm-on-start ann parallel/SIMD brute-force fallback below threshold
Vector / ANN cross-shard kNN merge, hybrid metadata pre-filter 🗺 ann single-shard today
Time-series store + time_bucket, ASOF join, gap-fill LOCF, OHLC, downsample, decay tsdb native redb columnar, no DataFusion
Time-series time-ops as unified planner ops (Op::Window) 🔶 tsdb functions ready; Op::Window is pass-through in the plan today
Blob / CAS content-addressed streaming store (redb-native) blob refcount mark-and-sweep GC; bounded RAM
Blob / CAS S3 / MinIO backend behind the same ChunkStore trait blob-s3 manifest/linkage byte-identical
Blob / CAS content-defined chunking 🗺 blob fixed 2 MiB chunks today
Key-value embedded in-process engine API over redb rows embedded EmbeddedEngine — no Tokio/socket/HMAC (KG-2.216)
Key-value generic get/put KV surface, SQLite-compatible wire 🗺 redb tables are graph-shaped; no KV/SQLite surface
Full-text Tantivy BM25 inverted index, RankText + reciprocal-rank fusion text composes in the unified planner
Unified planner Scan·Filter·Traverse·Rank·RankText·FuseRrf·Reason·SparqlBgp·Udf·ForeignScan·AsOf·Limit query+ each op feature-gated; see UQL
Unified planner Op::Window / Op::Foreign execution 🔶 query currently pass-through seams
UQL text DSL → wire::Plan (one parse, zero new exec path) (front-end always ships) dependency-free parser
UQL natural-language → query 🗺 only a reserved, rejected seam today
Durability redb-authoritative, commit-before-ack (kill -9-safe) redb folded into every tier
Distribution openraft replication + automatic failover raft cluster tier; off ⇒ byte-for-byte single-node
Distribution cross-shard 2PC (presumed-abort, crash-recoverable) raft classic blocking window; 3PC/non-blocking 🗺
Distribution multi-Raft groups (N-group resharding) 🔶 raft router scaffold; single DEFAULT_GROUP today
Federation remote engine / HTTP-JSON / external SQL (sqlx) as a ForeignScan federation(-sql) OFF by default; never in pi

Architecture at a glance

flowchart TB
    subgraph Clients["Clients"]
        AU["agent-utilities / graph-os"]
        PY["epistemic_graph Python client"]
        PG["psql / BI / ORM (pgwire)"]
        EMB["Embedded in-process caller (Pi/edge)"]
    end

    subgraph Engine["epistemic-graph-server (one Rust process)"]
        T["Transport: length-prefixed MessagePack over UDS / TCP, HMAC-SHA256"]
        SEC["Security: per-agent RLS + audit chain + encryption-at-rest"]
        PLAN["Unified RowSet planner (cost-reordered, cross-modal)"]
        CORE["GraphCore: petgraph + ledger + result cache"]

        subgraph Modalities["Modalities (feature-gated, one core)"]
            VEC["Vector ANN (eg-ann)"]
            SQL["SQL (eg-query / DataFusion)"]
            RDF["RDF / SPARQL / OWL (eg-rdf)"]
            TS["Time-series (eg-tsdb)"]
            TXT["Full-text (eg-text)"]
            BLOB["BLOB CAS (blob)"]
            WASM["WASM UDF (eg-wasm)"]
        end

        subgraph Durability["Durability and distribution"]
            REDB[("redb authoritative store")]
            RAFT["Raft replication + cross-shard 2PC (cluster)"]
            CDC["CDC / streaming / subscriptions"]
        end
    end

    AU --> T
    PY --> T
    PG --> SQL
    EMB --> CORE
    T --> SEC --> PLAN --> CORE
    CORE --> Modalities
    CORE --> REDB
    REDB <--> RAFT
    CORE --> CDC

A single cross-modal plan flows through one snapshot:

flowchart LR
    S["Scan / SparqlBgp / Reason<br/>(seed candidates)"] --> F["Filter<br/>(SQL predicates)"]
    F --> TR["Traverse<br/>(graph BFS)"]
    TR --> R["Rank / RankText<br/>(vector + BM25)"]
    R --> FU["FuseRrf<br/>(hybrid rank)"]
    FU --> A["AsOf<br/>(bi-temporal)"]
    A --> L["Limit"]

See docs/overview.md for the pipeline and docs/architecture/engine.md for the full architecture.


Deployment tiers and the prebuilt binaries

The same engine ships as a small family of prebuilt, size-optimized binaries (release-tiny profile). A Pi pulls a prebuilt wheel and never compiles. Full build/wheel recipes are in docs/deployment.md; the feature-composition map is in docs/architecture/tiers.md.

Binary Carries For
pi redb-authoritative + cypher + ann + rdf/sparql/owl + streaming + result-cache + cost — no DataFusion SQL, no Tantivy, no Raft Raspberry Pi / edge, ultra-lean
pi-max pi + tsdb + blob + security — all pure-Rust, still no C toolchain Pi "everything without a C compiler"
node pi + DataFusion SQL (query) + GraphQL + Tantivy text + owl-plan + wasm-udf + federation + finance/datascience single durable server
cluster node + Raft replication + pgwire + distributed compute + cross-shard 2PC multi-node HA / SQL clients
full every single-node feature, size-optimized (no raft/pgwire) workstation / one binary, every feature

Note: the lean pi tier carries SPARQL SELECT and OWL reasoning (via the Method::Owl* RPCs) but not the SQL-backed Op::Reason/Op::SparqlBgp planner ops — those need owl-plan, which pulls query (DataFusion) and lands in node. Every tier is redb-authoritative.


Three engine modes + the auto-bundle

agent-utilities reaches an engine through one resolver (EngineResolver, CONCEPT:OS-5.63) by a single precedence — no per-entrypoint code:

remote  ->  shared-local  ->  autostart

A configured remote (Docker on another host) is used as-is and never autostarts; a co-located engine already serving is reused; otherwise a detached, supervised engine is autostarted under a first-one-wins lock and reference-counted idle-shuts-down after its last client disconnects. Details + the decision flow: docs/engine-modes.md.

For the embedded/edge story, the embedded feature gives a SQLite/DuckDB-style in-process handle (EmbeddedEngine) over the same GraphCore + redb durable rows — no Tokio, no socket, no HMAC — the "100M agents, a local engine each" path.


Distribution & durability

  • redb-authoritative by default. A committed write is fsynced to redb before the client is acked (commit-before-ack); an acked write survives a hard crash. Eviction is read-through-safe.
  • In-engine Raft replication (cluster tier, raft). openraft replicates the authoritative redb store; the Raft log shares the one graph.redb (a log append + its graph mutation coalesce into one fsync). Leader failover is automatic. Off ⇒ the write path is byte-for-byte single-node.
  • Cross-shard 2PC. A transaction spanning multiple Raft groups commits atomically via presumed-abort two-phase commit, surviving coordinator/participant crashes. Multi-group routing/resharding is a scaffold today (single DEFAULT_GROUP); the durable machinery is in place.
  • Cross-modal ACID. A graph mutation + a vector upsert + a blob reference land in one redb WriteTransaction — all modalities commit together or none do.

Security & isolation

  • Auth is mandatory. Every RPC carries HMAC-SHA256(secret, request_id); the server refuses to start with an empty secret (--allow-insecure opts out, dev only). The pgwire surface adds SCRAM-SHA-256.
  • Per-agent Row-Level Security. Once any identity is registered, the read/plan-path GraphView is filtered to the rows the caller may see before any query surface (SQL/Cypher/SPARQL/GraphQL/unified) touches it. The result cache keys on the caller's RLS context.
  • Encryption-at-rest (security): redb durable value blobs are ChaCha20-Poly1305 AEAD-sealed (pure-Rust RustCrypto, no ring/openssl).
  • Hash-chained tamper-evident audit log over every durable mutation.

See docs/service_mode.md for the protocol, auth, and isolation policy.


Quickstart, per interface

Native client — out-of-process (the standard path)

from epistemic_graph import SyncEpistemicGraphClient

g = SyncEpistemicGraphClient()                    # connects/attaches to the UDS engine

g.nodes.add("AgentA", {"type": "coordinator"})
g.nodes.add("AgentB", {"type": "worker"})
g.edges.add("AgentA", "AgentB", {"weight": 1.5})
print("Order:", g.graph.topological_sort())

# OWL/RDFS forward chaining — materialises inferred edges/types in-graph
result = g.reasoning.reason(subclass_relations=[("Dog", "Animal")],
                            transitive_properties=["ancestor"])
print("Inferred:", result["inferred_count"], "triples")

Postgres wire (pgwire / cluster) — psql, BI tools, ORMs

# start the engine with the wire listener
EPISTEMIC_GRAPH_PGWIRE_ADDR=127.0.0.1:5433 \
  epistemic-graph-server --features cluster

# connect with any Postgres client
psql -h 127.0.0.1 -p 5433 -U agent -d epistemic
-- SELECT is full DataFusion: joins, aggregates, CTEs, window functions
SELECT n.id, n.properties->>'type' AS kind
FROM nodes n
WHERE n.properties->>'type' = 'worker';

-- DML on the graph node store (nodes table only, today)
INSERT INTO nodes (id, properties) VALUES ('AgentC', '{"type":"worker"}');
UPDATE nodes SET properties = '{"type":"idle"}' WHERE id = 'AgentC';
DELETE FROM nodes WHERE id = 'AgentC';

CREATE TABLE and arbitrary user tables are 🗺 roadmap; DML is the graph nodes table only today.

SPARQL (sparql)

g.rdf.add_triples([("ex:Dog", "rdfs:subClassOf", "ex:Animal")])

rows = g.rdf.sparql("""
  SELECT ?s ?o WHERE { ?s rdfs:subClassOf ?o }
""")                                              # SELECT supported today

ASK / CONSTRUCT / DESCRIBE / UPDATE and a /sparql HTTP endpoint are 🔶 in-progress — see the parity roadmap.

Embedded in-process (Pi / edge, embedded feature)

The EmbeddedEngine handle drives the same GraphCore + redb durable rows with no server, socket, or HMAC — open a persist dir and call core ops as plain methods (SQLite/DuckDB-style).

Batch, never per-element. Every out-of-process call is a serialize → socket → deserialize round trip, not a function call. Ship work as one batch op over data already in the graph; keep tight per-element math in-process. See AGENTS.md and docs/RUST_COMPUTE_GUIDE.md.


Ontology hosting & lifecycle

epistemic-graph is also an ontology server: you load OWL/RDFS as RDF, the engine maps it onto the property graph, and the EL⁺/RL reasoner materialises the closure (with confidence weights and Ebbinghaus time-decay). Classification, consistency checking, and incremental re-materialisation are all in-engine, and inferred members can seed a unified plan via REASON <Class>. See docs/interfaces/ontology.md for the load → reason → query → evolve lifecycle.


Documentation

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

epistemic_graph-2.0.0.tar.gz (1.4 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

epistemic_graph-2.0.0-py3-none-win_amd64.whl (29.6 MB view details)

Uploaded Python 3Windows x86-64

epistemic_graph-2.0.0-py3-none-manylinux_2_28_x86_64.whl (28.6 MB view details)

Uploaded Python 3manylinux: glibc 2.28+ x86-64

epistemic_graph-2.0.0-py3-none-manylinux_2_28_aarch64.whl (26.6 MB view details)

Uploaded Python 3manylinux: glibc 2.28+ ARM64

epistemic_graph-2.0.0-py3-none-macosx_11_0_arm64.whl (26.1 MB view details)

Uploaded Python 3macOS 11.0+ ARM64

File details

Details for the file epistemic_graph-2.0.0.tar.gz.

File metadata

  • Download URL: epistemic_graph-2.0.0.tar.gz
  • Upload date:
  • Size: 1.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for epistemic_graph-2.0.0.tar.gz
Algorithm Hash digest
SHA256 8b1e583740a0b52103b009eb21a78cf7371e72809a620265906ae48e3076a87c
MD5 ea5e1dcf764338ea9c7ded80365b07b3
BLAKE2b-256 a2909e83a5f2179b91239023b2e4372bf4c5b30bac51ae5cca01fc364e350962

See more details on using hashes here.

File details

Details for the file epistemic_graph-2.0.0-py3-none-win_amd64.whl.

File metadata

File hashes

Hashes for epistemic_graph-2.0.0-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 2d97c9210e10277a3c3443976826f931b14217ee71afcfe11e6a8bbdbfaf9d53
MD5 671363ed85147e6cb2edb94bd190100a
BLAKE2b-256 3984664031ab6c2d5aaa4b67641c824008d561f8851dbd3afef10fb5c10d845c

See more details on using hashes here.

File details

Details for the file epistemic_graph-2.0.0-py3-none-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for epistemic_graph-2.0.0-py3-none-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 539fcf4f4e006ee1a3ed6c1a2aa52c108bddaa24dac6c1f8c7f0dd414405c8fd
MD5 06f8c741f53242598c16b6744f85c7ff
BLAKE2b-256 87c50f48ca5284477d0cdb83bd1eb0a75710d6464462b9dbdb2df07f978f0c39

See more details on using hashes here.

File details

Details for the file epistemic_graph-2.0.0-py3-none-manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for epistemic_graph-2.0.0-py3-none-manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 532fcda024e0bbe3ff395acb190efaeb752f034f5debfd508938bbc940c57171
MD5 ede0657a85cf608873dfed6916ac1240
BLAKE2b-256 7e6a1c907f8edb73a1399fbb822f38959b2fffba3cee59ec0bce391025087284

See more details on using hashes here.

File details

Details for the file epistemic_graph-2.0.0-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for epistemic_graph-2.0.0-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 8d1144b1ad6e683bfb1151f33a528c5bd10d865cbc9da8f43105eb71bd0bcad7
MD5 8803aecf818b87eee29a12d24d7514f3
BLAKE2b-256 15d981ab8274109114d60489944b21d7295835032d4139e4a46865aa86e17827

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page