Rust-based epistemic graph engine for agent-utilities
Project description
epistemic-graph
One durable, Rust-native engine that is a drop-in substrate for graph · vector · SQL · SPARQL/RDF/OWL · time-series · blob · key-value
Every modality is a first-class citizen of one RowSet planner — from a Raspberry Pi to a replicated Raft cluster, from one core.
Honesty first. This README claims only what the code actually does today. Every row in the capability matrix is tagged ✅ supported · 🔶 in-progress · 🗺 roadmap, and the parity roadmap names exactly which gaps are being closed and in which order. If a doc and the code disagree, the code wins — file an issue.
Documentation — the full architecture, the tier/binary map, deployment recipes, the per-interface guides, and the concept registry live in the official documentation.
This is the compute & storage engine for
agent-utilities— a standalone Rust service reached out-of-process over MessagePack/UDS (no PyO3), embedded in-process, or spoken to over the Postgres wire protocol. Contributing? See CONTRIBUTING.md.
The thesis: one engine, every modality
A modern agent platform normally needs a graph database and a vector index and a SQL warehouse and a triple-store + reasoner and a time-series DB and a blob store and a full-text index — seven systems, seven copies of the data, six sync pipelines, and a brittle application layer that stitches results back together.
epistemic-graph collapses that rack into one durable engine with one unified query planner. Every
modality is a view over the same RowSet algebra, so a single plan can seed candidates from an OWL
inference or a SPARQL pattern, filter them with SQL, traverse the graph, re-rank by vector similarity
and BM25 text, fuse the two, and run a sandboxed WASM UDF — without ever leaving the engine or
marshalling rows back to Python.
It is durable by default: built with the redb feature (folded into every deployment tier), the
persist dir is the authoritative source of truth and an acked write survives kill -9
(commit-before-ack). It scales by configuration alone — the same binary family runs as an embedded
in-process library on a Pi, a single durable server, or a multi-node Raft cluster with cross-shard
transactions.
Drop-in positioning — and the honest parity status
| You run today | epistemic-graph as a drop-in | Current parity |
|---|---|---|
| Postgres (psql / BI / ORM) | pgwire server: SCRAM/trust auth, simple + extended protocol, pg_catalog/information_schema |
✅ read SQL + ✅ graph-table DML · 🗺 arbitrary tables + DDL |
| Stardog / GraphDB (RDF triple-store + reasoner) | RDF dataset over the property graph, SPARQL SELECT, OWL 2 EL⁺/RL reasoning | ✅ SELECT + ✅ reasoning · 🔶 ASK/CONSTRUCT/DESCRIBE/UPDATE + /sparql endpoint |
| Neo4j (property graph) | native petgraph core + Cypher MATCH…RETURN + graph algorithms |
✅ read traversal & algorithms · 🗺 Cypher writes |
| Pinecone / Milvus (vector DB) | native IVF-PQ + OPQ + SQ8 ANN, persistent, warm-on-start | ✅ |
| InfluxDB / TimescaleDB (time-series) | native redb TSDB: ASOF, gap-fill, time_bucket, OHLC, decay |
✅ primitives · 🔶 time-ops as planner ops |
| S3 / MinIO (blob) | content-addressed streaming CAS, redb-native or S3-backed | ✅ |
| SQLite / RocksDB (embedded KV) | EmbeddedEngine in-process handle over the same redb rows |
✅ embedded graph API · 🗺 generic KV / SQLite-wire surface |
The point is convergence, not a checkbox: the modalities share one snapshot, one ACID transaction, one security model, and one planner. See the full capability matrix below for the operation-by-operation truth.
Capability matrix
Legend: ✅ supported (implemented & tested) · 🔶 in-progress (partial or being added now) · 🗺 roadmap (designed, not built). Feature flags are the Cargo features that gate each surface; the tier table shows which prebuilt binary carries them.
| Interface | Operation | Status | Feature | Notes |
|---|---|---|---|---|
| SQL | SELECT (joins, aggregates, CTE, window, subquery) |
✅ | query |
DataFusion 43 over nodes + edges; real predicate pushdown (Inexact) |
| SQL | INSERT / UPDATE / DELETE |
✅ | query |
nodes table only, literal VALUES, single col = literal WHERE (KG-2.198) |
| SQL | Complex/compound WHERE, INSERT…SELECT, JOIN-in-DML |
🔶 | query |
KG-2.198 follow-up; explicitly errors today |
| SQL | Arbitrary user tables + DDL (CREATE/ALTER/DROP) |
🗺 | query |
errors unsupported statement; being added now |
| Postgres wire | listener, simple + extended/prepared protocol | ✅ | pgwire |
EPISTEMIC_GRAPH_PGWIRE_ADDR; also pulled in by cluster |
| Postgres wire | SCRAM-SHA-256 / trust auth, pg_catalog introspection |
✅ | pgwire |
KG-2.202 / KG-2.201; pg user → engine ACL actor |
| SPARQL | SELECT (BGP, paths, FILTER subset, OPTIONAL, UNION, GROUP/agg, BIND, DISTINCT, SLICE) |
✅ | sparql |
spargebra parser compiled to LPG scans |
| SPARQL | ASK / CONSTRUCT / DESCRIBE |
🔶 | sparql |
errors supports SELECT only; being added now |
| SPARQL | UPDATE (INSERT/DELETE DATA) |
🗺 | sparql |
absent in eg-rdf |
| SPARQL | /sparql HTTP endpoint |
🗺 | — | today exposed as binary RPC Method::Sparql |
| SPARQL | true named graphs, SPO/POS index, regex/arith FILTER | 🔶 | sparql |
single default graph; naive full-scan; FILTER is a subset |
| Cypher | MATCH … WHERE … RETURN … LIMIT (var-length single hop) |
✅ | cypher |
read-only over one snapshot; WHERE is AND-only |
| Cypher | writes (CREATE/MERGE/SET/DELETE), ORDER BY/WITH/OR |
🗺 | cypher |
not in grammar |
| GraphQL | read queries (scan + BFS, schema-from-graph, aliases, first/limit, filters) |
✅ | graphql |
byte-equal to Cypher path |
| GraphQL | mutations / subscriptions / fragments / variables | 🗺 | graphql |
explicitly rejected at parse |
| OWL | EL⁺ + RL forward-chaining materialization & classification | ✅ | owl |
pure-Rust; consistency + incremental + justifications |
| OWL | confidence-weighting + Ebbinghaus time-decay | ✅ | owl |
KG-2.236; per-axiom eg:confidence, fact decay |
| OWL | query-time Op::Reason (reasoner seeds a RowSet) |
✅ | owl-plan |
distributed/cross-shard union supported |
| OWL | OWL-DL (tableau, cardinality, allValuesFrom), SWRL user rules |
🗺 | — | out of the EL+RL envelope by design |
| Vector / ANN | IVF-PQ + OPQ + SQ8-refine, persistent (reopen w/o rebuild), warm-on-start | ✅ | ann |
parallel/SIMD brute-force fallback below threshold |
| Vector / ANN | cross-shard kNN merge, hybrid metadata pre-filter | 🗺 | ann |
single-shard today |
| Time-series | store + time_bucket, ASOF join, gap-fill LOCF, OHLC, downsample, decay |
✅ | tsdb |
native redb columnar, no DataFusion |
| Time-series | time-ops as unified planner ops (Op::Window) |
🔶 | tsdb |
functions ready; Op::Window is pass-through in the plan today |
| Blob / CAS | content-addressed streaming store (redb-native) | ✅ | blob |
refcount mark-and-sweep GC; bounded RAM |
| Blob / CAS | S3 / MinIO backend behind the same ChunkStore trait |
✅ | blob-s3 |
manifest/linkage byte-identical |
| Blob / CAS | content-defined chunking | 🗺 | blob |
fixed 2 MiB chunks today |
| Key-value | embedded in-process engine API over redb rows | ✅ | embedded |
EmbeddedEngine — no Tokio/socket/HMAC (KG-2.216) |
| Key-value | generic get/put KV surface, SQLite-compatible wire |
🗺 | — | redb tables are graph-shaped; no KV/SQLite surface |
| Full-text | Tantivy BM25 inverted index, RankText + reciprocal-rank fusion |
✅ | text |
composes in the unified planner |
| Unified planner | Scan·Filter·Traverse·Rank·RankText·FuseRrf·Reason·SparqlBgp·Udf·ForeignScan·AsOf·Limit |
✅ | query+ |
each op feature-gated; see UQL |
| Unified planner | Op::Window / Op::Foreign execution |
🔶 | query |
currently pass-through seams |
| UQL | text DSL → wire::Plan (one parse, zero new exec path) |
✅ | (front-end always ships) | dependency-free parser |
| UQL | natural-language → query | 🗺 | — | only a reserved, rejected seam today |
| Durability | redb-authoritative, commit-before-ack (kill -9-safe) |
✅ | redb |
folded into every tier |
| Distribution | openraft replication + automatic failover | ✅ | raft |
cluster tier; off ⇒ byte-for-byte single-node |
| Distribution | cross-shard 2PC (presumed-abort, crash-recoverable) | ✅ | raft |
classic blocking window; 3PC/non-blocking 🗺 |
| Distribution | multi-Raft groups (N-group resharding) | 🔶 | raft |
router scaffold; single DEFAULT_GROUP today |
| Federation | remote engine / HTTP-JSON / external SQL (sqlx) as a ForeignScan |
✅ | federation(-sql) |
OFF by default; never in pi |
Architecture at a glance
flowchart TB
subgraph Clients["Clients"]
AU["agent-utilities / graph-os"]
PY["epistemic_graph Python client"]
PG["psql / BI / ORM (pgwire)"]
EMB["Embedded in-process caller (Pi/edge)"]
end
subgraph Engine["epistemic-graph-server (one Rust process)"]
T["Transport: length-prefixed MessagePack over UDS / TCP, HMAC-SHA256"]
SEC["Security: per-agent RLS + audit chain + encryption-at-rest"]
PLAN["Unified RowSet planner (cost-reordered, cross-modal)"]
CORE["GraphCore: petgraph + ledger + result cache"]
subgraph Modalities["Modalities (feature-gated, one core)"]
VEC["Vector ANN (eg-ann)"]
SQL["SQL (eg-query / DataFusion)"]
RDF["RDF / SPARQL / OWL (eg-rdf)"]
TS["Time-series (eg-tsdb)"]
TXT["Full-text (eg-text)"]
BLOB["BLOB CAS (blob)"]
WASM["WASM UDF (eg-wasm)"]
end
subgraph Durability["Durability and distribution"]
REDB[("redb authoritative store")]
RAFT["Raft replication + cross-shard 2PC (cluster)"]
CDC["CDC / streaming / subscriptions"]
end
end
AU --> T
PY --> T
PG --> SQL
EMB --> CORE
T --> SEC --> PLAN --> CORE
CORE --> Modalities
CORE --> REDB
REDB <--> RAFT
CORE --> CDC
A single cross-modal plan flows through one snapshot:
flowchart LR
S["Scan / SparqlBgp / Reason<br/>(seed candidates)"] --> F["Filter<br/>(SQL predicates)"]
F --> TR["Traverse<br/>(graph BFS)"]
TR --> R["Rank / RankText<br/>(vector + BM25)"]
R --> FU["FuseRrf<br/>(hybrid rank)"]
FU --> A["AsOf<br/>(bi-temporal)"]
A --> L["Limit"]
See docs/overview.md for the pipeline and docs/architecture/engine.md for the full architecture.
Deployment tiers and the prebuilt binaries
The same engine ships as a small family of prebuilt, size-optimized binaries (release-tiny profile).
A Pi pulls a prebuilt wheel and never compiles. Full build/wheel recipes are in
docs/deployment.md; the feature-composition map is in
docs/architecture/tiers.md.
| Binary | Carries | For |
|---|---|---|
| pi | redb-authoritative + cypher + ann + rdf/sparql/owl + streaming + result-cache + cost — no DataFusion SQL, no Tantivy, no Raft | Raspberry Pi / edge, ultra-lean |
| pi-max | pi + tsdb + blob + security — all pure-Rust, still no C toolchain | Pi "everything without a C compiler" |
| node | pi + DataFusion SQL (query) + GraphQL + Tantivy text + owl-plan + wasm-udf + federation + finance/datascience |
single durable server |
| cluster | node + Raft replication + pgwire + distributed compute + cross-shard 2PC | multi-node HA / SQL clients |
| full | every single-node feature, size-optimized (no raft/pgwire) | workstation / one binary, every feature |
Note: the lean pi tier carries SPARQL
SELECTand OWL reasoning (via theMethod::Owl*RPCs) but not the SQL-backedOp::Reason/Op::SparqlBgpplanner ops — those needowl-plan, which pullsquery(DataFusion) and lands in node. Every tier is redb-authoritative.
Three engine modes + the auto-bundle
agent-utilities reaches an engine through one resolver (EngineResolver, CONCEPT:OS-5.63) by a
single precedence — no per-entrypoint code:
remote -> shared-local -> autostart
A configured remote (Docker on another host) is used as-is and never autostarts; a co-located engine already serving is reused; otherwise a detached, supervised engine is autostarted under a first-one-wins lock and reference-counted idle-shuts-down after its last client disconnects. Details + the decision flow: docs/engine-modes.md.
For the embedded/edge story, the embedded feature gives a SQLite/DuckDB-style in-process handle
(EmbeddedEngine) over the same GraphCore + redb durable rows — no Tokio, no socket, no HMAC — the
"100M agents, a local engine each" path.
Distribution & durability
- redb-authoritative by default. A committed write is fsynced to redb before the client is acked (commit-before-ack); an acked write survives a hard crash. Eviction is read-through-safe.
- In-engine Raft replication (cluster tier,
raft).openraftreplicates the authoritative redb store; the Raft log shares the onegraph.redb(a log append + its graph mutation coalesce into one fsync). Leader failover is automatic. Off ⇒ the write path is byte-for-byte single-node. - Cross-shard 2PC. A transaction spanning multiple Raft groups commits atomically via presumed-abort
two-phase commit, surviving coordinator/participant crashes. Multi-group routing/resharding is a
scaffold today (single
DEFAULT_GROUP); the durable machinery is in place. - Cross-modal ACID. A graph mutation + a vector upsert + a blob reference land in one redb
WriteTransaction— all modalities commit together or none do.
Security & isolation
- Auth is mandatory. Every RPC carries
HMAC-SHA256(secret, request_id); the server refuses to start with an empty secret (--allow-insecureopts out, dev only). The pgwire surface adds SCRAM-SHA-256. - Per-agent Row-Level Security. Once any identity is registered, the read/plan-path
GraphViewis filtered to the rows the caller may see before any query surface (SQL/Cypher/SPARQL/GraphQL/unified) touches it. The result cache keys on the caller's RLS context. - Encryption-at-rest (
security): redb durable value blobs are ChaCha20-Poly1305 AEAD-sealed (pure-Rust RustCrypto, no ring/openssl). - Hash-chained tamper-evident audit log over every durable mutation.
See docs/service_mode.md for the protocol, auth, and isolation policy.
Quickstart, per interface
Native client — out-of-process (the standard path)
from epistemic_graph import SyncEpistemicGraphClient
g = SyncEpistemicGraphClient() # connects/attaches to the UDS engine
g.nodes.add("AgentA", {"type": "coordinator"})
g.nodes.add("AgentB", {"type": "worker"})
g.edges.add("AgentA", "AgentB", {"weight": 1.5})
print("Order:", g.graph.topological_sort())
# OWL/RDFS forward chaining — materialises inferred edges/types in-graph
result = g.reasoning.reason(subclass_relations=[("Dog", "Animal")],
transitive_properties=["ancestor"])
print("Inferred:", result["inferred_count"], "triples")
Postgres wire (pgwire / cluster) — psql, BI tools, ORMs
# start the engine with the wire listener
EPISTEMIC_GRAPH_PGWIRE_ADDR=127.0.0.1:5433 \
epistemic-graph-server --features cluster
# connect with any Postgres client
psql -h 127.0.0.1 -p 5433 -U agent -d epistemic
-- SELECT is full DataFusion: joins, aggregates, CTEs, window functions
SELECT n.id, n.properties->>'type' AS kind
FROM nodes n
WHERE n.properties->>'type' = 'worker';
-- DML on the graph node store (nodes table only, today)
INSERT INTO nodes (id, properties) VALUES ('AgentC', '{"type":"worker"}');
UPDATE nodes SET properties = '{"type":"idle"}' WHERE id = 'AgentC';
DELETE FROM nodes WHERE id = 'AgentC';
CREATE TABLEand arbitrary user tables are 🗺 roadmap; DML is the graphnodestable only today.
SPARQL (sparql)
g.rdf.add_triples([("ex:Dog", "rdfs:subClassOf", "ex:Animal")])
rows = g.rdf.sparql("""
SELECT ?s ?o WHERE { ?s rdfs:subClassOf ?o }
""") # SELECT supported today
ASK/CONSTRUCT/DESCRIBE/UPDATEand a/sparqlHTTP endpoint are 🔶 in-progress — see the parity roadmap.
Embedded in-process (Pi / edge, embedded feature)
The EmbeddedEngine handle drives the same GraphCore + redb durable rows with no server, socket, or
HMAC — open a persist dir and call core ops as plain methods (SQLite/DuckDB-style).
Batch, never per-element. Every out-of-process call is a serialize → socket → deserialize round trip, not a function call. Ship work as one batch op over data already in the graph; keep tight per-element math in-process. See AGENTS.md and docs/RUST_COMPUTE_GUIDE.md.
Ontology hosting & lifecycle
epistemic-graph is also an ontology server: you load OWL/RDFS as RDF, the engine maps it onto the
property graph, and the EL⁺/RL reasoner materialises the closure (with confidence weights and Ebbinghaus
time-decay). Classification, consistency checking, and incremental re-materialisation are all
in-engine, and inferred members can seed a unified plan via REASON <Class>. See
docs/interfaces/ontology.md for the load → reason → query → evolve
lifecycle.
Documentation
- Capabilities & parity matrix — the operation-by-operation truth table.
- Universal-DB parity roadmap — every gap being closed, with status.
- Technical Overview · Master-of-all engine.
- Per-interface guides: SQL · SPARQL · Cypher · GraphQL · Vector · Time-series · KV & Blob · Ontology lifecycle.
- UQL & the unified planner · Tiers & binaries · Deployment · Engine modes · Service Mode.
License
MIT — see LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file epistemic_graph-2.0.0.tar.gz.
File metadata
- Download URL: epistemic_graph-2.0.0.tar.gz
- Upload date:
- Size: 1.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8b1e583740a0b52103b009eb21a78cf7371e72809a620265906ae48e3076a87c
|
|
| MD5 |
ea5e1dcf764338ea9c7ded80365b07b3
|
|
| BLAKE2b-256 |
a2909e83a5f2179b91239023b2e4372bf4c5b30bac51ae5cca01fc364e350962
|
File details
Details for the file epistemic_graph-2.0.0-py3-none-win_amd64.whl.
File metadata
- Download URL: epistemic_graph-2.0.0-py3-none-win_amd64.whl
- Upload date:
- Size: 29.6 MB
- Tags: Python 3, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2d97c9210e10277a3c3443976826f931b14217ee71afcfe11e6a8bbdbfaf9d53
|
|
| MD5 |
671363ed85147e6cb2edb94bd190100a
|
|
| BLAKE2b-256 |
3984664031ab6c2d5aaa4b67641c824008d561f8851dbd3afef10fb5c10d845c
|
File details
Details for the file epistemic_graph-2.0.0-py3-none-manylinux_2_28_x86_64.whl.
File metadata
- Download URL: epistemic_graph-2.0.0-py3-none-manylinux_2_28_x86_64.whl
- Upload date:
- Size: 28.6 MB
- Tags: Python 3, manylinux: glibc 2.28+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
539fcf4f4e006ee1a3ed6c1a2aa52c108bddaa24dac6c1f8c7f0dd414405c8fd
|
|
| MD5 |
06f8c741f53242598c16b6744f85c7ff
|
|
| BLAKE2b-256 |
87c50f48ca5284477d0cdb83bd1eb0a75710d6464462b9dbdb2df07f978f0c39
|
File details
Details for the file epistemic_graph-2.0.0-py3-none-manylinux_2_28_aarch64.whl.
File metadata
- Download URL: epistemic_graph-2.0.0-py3-none-manylinux_2_28_aarch64.whl
- Upload date:
- Size: 26.6 MB
- Tags: Python 3, manylinux: glibc 2.28+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
532fcda024e0bbe3ff395acb190efaeb752f034f5debfd508938bbc940c57171
|
|
| MD5 |
ede0657a85cf608873dfed6916ac1240
|
|
| BLAKE2b-256 |
7e6a1c907f8edb73a1399fbb822f38959b2fffba3cee59ec0bce391025087284
|
File details
Details for the file epistemic_graph-2.0.0-py3-none-macosx_11_0_arm64.whl.
File metadata
- Download URL: epistemic_graph-2.0.0-py3-none-macosx_11_0_arm64.whl
- Upload date:
- Size: 26.1 MB
- Tags: Python 3, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8d1144b1ad6e683bfb1151f33a528c5bd10d865cbc9da8f43105eb71bd0bcad7
|
|
| MD5 |
8803aecf818b87eee29a12d24d7514f3
|
|
| BLAKE2b-256 |
15d981ab8274109114d60489944b21d7295835032d4139e4a46865aa86e17827
|