Document knowledge graphs — extract structured data from documents (Confluence, PDFs, CSV, JSON, Excel) into Neptune
Project description
Document Graph
Document knowledge graphs — extract structured data from documents (Confluence, PDFs, CSV, JSON, Excel) into Neptune.
This package depends on AWS GraphRAG Toolkit (graphrag-toolkit-lexical-graph) for graph storage, vector indexing, and retrieval.
Quick Start
Install
pip install document-graph # core
pip install document-graph[graphrag] # with lexical-graph integration
Python API Example
from graphrag_toolkit.document_graph import PipelineExecutor, Node, Edge
from graphrag_toolkit.document_graph.graph_build.cypher_builder import CypherBuilder
# Create typed nodes from structured data
node = Node(
id="doc-001",
labels=["Document"],
properties={"title": "Q4 Report", "source": "confluence", "tenant_id": "acme"}
)
# Build Cypher for Neptune ingestion
cypher, params = CypherBuilder.node_to_cypher(node, tenant_id="acme")
# Execute against Neptune via graphrag-toolkit GraphStore
graph_store.execute_query(cypher, params)
Schema-Driven Pipeline
from graphrag_toolkit.document_graph.schema.providers.csv_schema_provider import CSVSchemaProvider
from graphrag_toolkit.document_graph.schema.providers.schema_provider_config import SchemaProviderConfig
# Auto-discover schema from CSV
config = SchemaProviderConfig(type="csv", connection_config={"path": "data/employees.csv"})
provider = CSVSchemaProvider(config)
schema = provider.load_schema()
# Transform and load
from graphrag_toolkit.document_graph.transform.graph_transformers.row_to_node import RowToNodeTransformer
from graphrag_toolkit.document_graph.transform.transformer_provider_config import TransformerProviderConfig
transformer = RowToNodeTransformer(TransformerProviderConfig(name="r2n", args={"type": "Employee"}))
nodes = transformer.transform(records)
Hybrid Search (Document Graph + Lexical Graph)
from graphrag_toolkit.lexical_graph import LexicalGraphIndex, LexicalGraphQueryEngine
from graphrag_toolkit.lexical_graph.storage import GraphStoreFactory, VectorStoreFactory
from llama_index.core.schema import Document
# Write document-graph nodes, then index into lexical-graph for semantic search
graph_store = GraphStoreFactory.for_graph_store("neptune-db://endpoint:8182").__enter__()
vector_store = VectorStoreFactory.for_vector_store("aoss://endpoint")
graph_index = LexicalGraphIndex(graph_store, vector_store)
graph_index.extract_and_build(docs, show_progress=True)
# Semantic query across both structured and unstructured data
query_engine = LexicalGraphQueryEngine.for_traversal_based_search(graph_store, vector_store)
results = query_engine.retrieve("Who are the senior engineers?")
Package Structure
src/graphrag_toolkit/document_graph/
├── __init__.py # Public API: PipelineExecutor, Node, Edge, models
├── config.py # Configuration
├── errors.py # Custom exceptions
├── model.py # NodeModel, EdgeModel
├── model_elements.py # Node, Edge primitives
├── pipeline_executor.py # Orchestrates ingest → transform → build
├── schema/ # ETL schema model, providers, discovery
│ ├── providers/ # CSV, JSON, S3, Static, File, Glue
│ └── discovery/ # Auto-infer schema from data files
├── ingest/ # Data ingestion (column, field, row processors)
├── transform/ # 20+ transformers
│ ├── normalizers/ # Whitespace, nulls, case, enum, timestamp
│ ├── field_transformers/ # JSON flattener, UUID gen, regex clean
│ ├── document_transformers/ # JSON to rows, text chunker, PII redactor
│ ├── filter_transformers/ # Row filter, column pruner
│ ├── graph_transformers/ # Row to node, infer edges
│ └── truncators/ # Length, field count, token limits
├── graph_build/ # Cypher generation with tenant-scoped labels
│ └── constructors/ # Node/edge Cypher builders
├── query/ # DocumentGraphQueryEngine
├── pipeline/ # Extract (CSV, Excel, JSON, Parquet) and load
└── plugins/ # Plugin system for extensibility
Integration
With AWS GraphRAG Toolkit
Document Graph extends the AWS GraphRAG Toolkit architecture:
┌─────────────────────────────────────────────────────┐
│ document-graph (this package) │
│ Structured ETL: CSV, Excel, JSON, PDF → typed nodes│
├─────────────────────────────────────────────────────┤
│ graphrag-toolkit-lexical-graph (foundation) │
│ GraphStore, VectorStore, Neptune/AOSS writers │
│ LexicalGraphIndex, entity resolution, retrieval │
└─────────────────────────────────────────────────────┘
With Neptune & OpenSearch Serverless
Deploy infrastructure via the graphrag-toolkit CloudFormation templates which provisions:
- Amazon Neptune (graph database)
- Amazon OpenSearch Serverless (vector search)
- SageMaker notebook instance
Multi-Tenancy
All operations use tenant-scoped labels for complete data isolation:
# Tenant A's data is completely isolated from Tenant B
node_to_cypher(node, tenant_id="acme_corp") # → MERGE (n:`__User__acme_corp__` ...)
node_to_cypher(node, tenant_id="beta_inc") # → MERGE (n:`__User__beta_inc__` ...)
Requirements
- Python >= 3.11
pydantic >= 2.0pandas >= 2.0boto3 >= 1.26- Optional:
graphrag-toolkit-lexical-graph >= 3.18.0(for hybrid search)
License
MIT — see LICENSE for details.
See NOTICE for third-party acknowledgments.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file graphrag_document_graph-0.1.0.tar.gz.
File metadata
- Download URL: graphrag_document_graph-0.1.0.tar.gz
- Upload date:
- Size: 230.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f9e285249f6aee2b07a87ae047ef800c4075097e48b0cee9d07246edee6bd2f1
|
|
| MD5 |
2c44100304ad59f71b4ca81155f77503
|
|
| BLAKE2b-256 |
c7a03698fd73840ef8922c635d05c82411a262983f46684b4f653433baf1d40e
|
File details
Details for the file graphrag_document_graph-0.1.0-py3-none-any.whl.
File metadata
- Download URL: graphrag_document_graph-0.1.0-py3-none-any.whl
- Upload date:
- Size: 255.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5bc1f36a324fad0a975209561d2c06dc562038f723b14e2862c4a05ca5253b7d
|
|
| MD5 |
0217a5bfcdd6039055ea6d08526cd4d2
|
|
| BLAKE2b-256 |
56402e64a7d2f15c0106118c13f26171454d384717754cb8ab0658440fb8076e
|