unstructured2graph
Convert unstructured documents into knowledge graphs within Memgraph.
Overview
unstructured2graph enables you to transform any unstructured data (PDFs, URLs, documents) into a graph database, powering Graph Retrieval-Augmented Generation (GraphRAG) applications. It combines:
- Unstructured - Parse and chunk diverse document formats
- LightRAG - Extract entities and relationships using LLMs
- Memgraph - Store and query your knowledge graph
Installation
Install from source:
git clone https://github.com/memgraph/ai-toolkit.git
cd ai-toolkit/unstructured2graph
pip install -e .
For full document support (PDF, DOCX, etc.):
pip install -e ".[all-docs]"
Quick Start
import asyncio
from memgraph_toolbox.api.memgraph import Memgraph
from lightrag_memgraph import MemgraphLightRAGWrapper
from unstructured2graph import from_unstructured, create_property_index
async def main():
memgraph = Memgraph(user_agent="unstructured2graph")
create_property_index(memgraph, "Chunk", "hash")
lightrag = MemgraphLightRAGWrapper()
await lightrag.initialize(working_dir="./lightrag_storage")
# Ingest documents from URLs or local files
await from_unstructured(
sources=["https://example.com/doc.pdf", "./local_file.md"],
memgraph=memgraph,
lightrag_wrapper=lightrag,
link_chunks=True, # Create NEXT relationships between chunks
)
await lightrag.afinalize()
asyncio.run(main())
Persistence:
MemgraphLightRAGWrappernow persists LightRAG's full working state into Memgraph by default — the entity/relationship graph plus the key/value store, vector store, and document-status store. Theworking_dirargument is still accepted (and used as a fallback location for any store not backed by Memgraph), but with the default settings the JSON stores are no longer written there. See the lightrag-memgraph README for the label/index schema and opt-out flags.
Key Features
| Feature | Description |
|---|---|
| Multi-format parsing | PDFs, URLs, HTML, Markdown, DOCX, and more via Unstructured |
| Automatic chunking | Smart document chunking with configurable options |
| Entity extraction | LLM-powered entity and relationship extraction via LightRAG |
| Vector search | Built-in support for embedding generation and vector indices |
| GraphRAG queries | Combine vector search with graph traversal for enhanced retrieval |
API Reference
Document Processing
parse_source(source, partition_kwargs)- Parse a single file or URL into chunksparse_text(text, partition_kwargs)- Chunk a raw in-memory string (no file/URL involved)make_chunks(sources, partition_kwargs)- Process multiple sources intoChunkedDocumentobjectsfrom_unstructured(sources, memgraph, lightrag_wrapper, ...)- Full ingestion pipeline for files/URLs; returns one Chunk group per sourcefrom_texts(texts, memgraph, lightrag_wrapper, ...)- Full ingestion pipeline for raw strings; returns one Chunk group per input text
Graph Operations
create_nodes_from_list(memgraph, nodes, label, batch_size)- Batch insert nodesconnect_chunks_to_entities(memgraph, chunk_label, entity_label)- Link entities to source chunkslink_nodes_in_order(memgraph, ...)- Create sequential relationships between chunkscreate_vector_search_index(memgraph, label, property)- Create vector index for similarity searchcompute_embeddings(memgraph, label)- Generate embeddings for nodes
Documentation
For detailed usage examples and getting started guides, check out the official documentation:
👉 unstructured2graph Documentation
Requirements
- Python 3.10+
- Memgraph database instance
LLM API Key
This library uses LightRAG for entity and relationship extraction, which requires an LLM API key. Set your OpenAI API key as an environment variable:
export OPENAI_API_KEY="your-api-key"
Metadata
Release files for unstructured2graph 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| unstructured2graph-0.4.0.tar.gz | 17.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| unstructured2graph-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 27.6 kB
Release files / unstructured2graph-0.4.0.tar.gz
| Download URL | unstructured2graph-0.4.0.tar.gz |
|---|---|
| Size | 17.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0a7d1912e22fb233f63f258356ac1b825e15af1d435dceffc7fecb42870221ab
|
|
BLAKE2b-256 checksum How to use checksums |
a6a787324f728ceb9cf53ff175109fc9225e0603996d3c4040cd4642bbb68c0c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / unstructured2graph-0.4.0-py3-none-any.whl
| Download URL | unstructured2graph-0.4.0-py3-none-any.whl |
|---|---|
| Size | 10.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
25d96d3b162e0b4e5649825cf494db7d8898b118cf55d5410e0a41c58f38b4f2
|
|
BLAKE2b-256 checksum How to use checksums |
763d4a47d8d5c14a454c2ff6643fc94e20c8237705cf8fb874e2f80ef8b7826c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|