Document Graph
Document knowledge graphs — extract structured data from documents (Confluence, PDFs, CSV, JSON, Excel) into Neptune.
This package depends on AWS GraphRAG Toolkit (graphrag-toolkit-lexical-graph) for graph storage, vector indexing, and retrieval.
Installation
pip install graphrag-document-graph
Dependencies
graphrag-toolkit-lexical-graph>=3.18.0— AWS GraphRAG Toolkit foundation (graph storage, vector indexing, Neptune/AOSS writers)
This package does not depend on bona or graphrag-codeproperty-graph.
Dependency Chain
graphrag-document-graph
└── graphrag-toolkit-lexical-graph>=3.18.0 (AWS foundation)
├── Neptune graph storage
├── OpenSearch Serverless vector indexing
└── Entity resolution & retrieval
Quick Start
Python API Example
from graphrag_toolkit.document_graph import PipelineExecutor, Node, Edge
from graphrag_toolkit.document_graph.graph_build.cypher_builder import CypherBuilder
# Create typed nodes from structured data
node = Node(
id="doc-001",
labels=["Document"],
properties={"title": "Q4 Report", "source": "confluence", "tenant_id": "acme"}
)
# Build Cypher for Neptune ingestion
cypher, params = CypherBuilder.node_to_cypher(node, tenant_id="acme")
# Execute against Neptune via graphrag-toolkit GraphStore
graph_store.execute_query(cypher, params)
Schema-Driven Pipeline
from graphrag_toolkit.document_graph.schema.providers.csv_schema_provider import CSVSchemaProvider
from graphrag_toolkit.document_graph.schema.providers.schema_provider_config import SchemaProviderConfig
# Auto-discover schema from CSV
config = SchemaProviderConfig(type="csv", connection_config={"path": "data/employees.csv"})
provider = CSVSchemaProvider(config)
schema = provider.load_schema()
# Transform and load
from graphrag_toolkit.document_graph.transform.graph_transformers.row_to_node import RowToNodeTransformer
from graphrag_toolkit.document_graph.transform.transformer_provider_config import TransformerProviderConfig
transformer = RowToNodeTransformer(TransformerProviderConfig(name="r2n", args={"type": "Employee"}))
nodes = transformer.transform(records)
Hybrid Search (Document Graph + Lexical Graph)
from graphrag_toolkit.lexical_graph import LexicalGraphIndex, LexicalGraphQueryEngine
from graphrag_toolkit.lexical_graph.storage import GraphStoreFactory, VectorStoreFactory
# Write document-graph nodes, then index into lexical-graph for semantic search
graph_store = GraphStoreFactory.for_graph_store("neptune-db://endpoint:8182").__enter__()
vector_store = VectorStoreFactory.for_vector_store("aoss://endpoint")
graph_index = LexicalGraphIndex(graph_store, vector_store)
graph_index.extract_and_build(docs, show_progress=True)
# Semantic query across both structured and unstructured data
query_engine = LexicalGraphQueryEngine.for_traversal_based_search(graph_store, vector_store)
results = query_engine.retrieve("Who are the senior engineers?")
Package Structure
src/graphrag_toolkit/document_graph/
├── __init__.py # Public API: PipelineExecutor, Node, Edge, models
├── config.py # Configuration
├── errors.py # Custom exceptions
├── model.py # NodeModel, EdgeModel
├── model_elements.py # Node, Edge primitives
├── pipeline_executor.py # Orchestrates ingest → transform → build
├── schema/ # ETL schema model, providers, discovery
│ ├── providers/ # CSV, JSON, S3, Static, File, Glue
│ └── discovery/ # Auto-infer schema from data files
├── ingest/ # Data ingestion (column, field, row processors)
├── transform/ # 20+ transformers
│ ├── normalizers/ # Whitespace, nulls, case, enum, timestamp
│ ├── field_transformers/ # JSON flattener, UUID gen, regex clean
│ ├── document_transformers/ # JSON to rows, text chunker, PII redactor
│ ├── filter_transformers/ # Row filter, column pruner
│ ├── graph_transformers/ # Row to node, infer edges
│ └── truncators/ # Length, field count, token limits
├── graph_build/ # Cypher generation with tenant-scoped labels
│ └── constructors/ # Node/edge Cypher builders
├── query/ # DocumentGraphQueryEngine
├── pipeline/ # Extract (CSV, Excel, JSON, Parquet) and load
└── plugins/ # Plugin system for extensibility
Integration with AWS GraphRAG Toolkit
Document Graph extends the AWS GraphRAG Toolkit architecture:
┌─────────────────────────────────────────────────────┐
│ document-graph (this package) │
│ Structured ETL: CSV, Excel, JSON, PDF → typed nodes│
├─────────────────────────────────────────────────────┤
│ graphrag-toolkit-lexical-graph (foundation) │
│ GraphStore, VectorStore, Neptune/AOSS writers │
│ LexicalGraphIndex, entity resolution, retrieval │
└─────────────────────────────────────────────────────┘
Multi-Tenancy
All operations use tenant-scoped labels for complete data isolation:
node_to_cypher(node, tenant_id="acme_corp") # → MERGE (n:`__User__acme_corp__` ...)
node_to_cypher(node, tenant_id="beta_inc") # → MERGE (n:`__User__beta_inc__` ...)
Contributing
See CONTRIBUTING.md for development setup, testing, and PR guidelines.
License
MIT — see LICENSE for details.
See NOTICE for third-party acknowledgments.
Release files for graphrag-document-graph 3.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| graphrag_document_graph-3.1.2.tar.gz | 699.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| graphrag_document_graph-3.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 961.5 kB
Release files / graphrag_document_graph-3.1.2.tar.gz
| Download URL | graphrag_document_graph-3.1.2.tar.gz |
|---|---|
| Size | 699.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3eb9320314002858cb26b7aca0293994483646090f817d4077a0251adeea940e
|
|
BLAKE2b-256 checksum How to use checksums |
d14f5f8d4c32e301ac8fa0704354bedffcc46b2179a39c5064a74fa4797d6a7b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / graphrag_document_graph-3.1.2-py3-none-any.whl
| Download URL | graphrag_document_graph-3.1.2-py3-none-any.whl |
|---|---|
| Size | 262.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ac6f25a68ebdf5f9d368329e254a5eab161247cef48b575c612a42dc0a5813ca
|
|
BLAKE2b-256 checksum How to use checksums |
4979ab5e087da714f0f0caae65ca46f4912f5e4e5a3a36c253f505aff7cb0866
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|