A simple graph-based retriever using Neo4j and Qdrant
Project description
Simple Graph Retriever
A Python SDK for indexing a graph database (Neo4j or FalkorDB) into a vector database (Qdrant) and performing retrieval-augmented generation (RAG) tasks against it.
This SDK handles the heavy lifting of graph processing, including:
- Community Detection: Uses the Leiden algorithm to identify communities within your graph structure.
- Chunking: Breaks down nodes and their local context into "chunks" suitable for embedding.
- Vector Indexing: Embeds and indexes both community summaries and individual chunks into Qdrant.
- Graph Retrieval: Retrieves relevant subgraphs based on a natural language query, with fine-grained control over the expansion process.
Table of Contents
- Simple Graph Retriever
1. Installation
To install the SDK, clone this repository and use pip. You can choose which database driver to install using optional dependencies.
Install with Neo4j support:
pip install simple_graph_retriever[neo4j]==1.0.0-rc.17
Install with FalkorDB support:
pip install simple_graph_retriever[falkordb]==1.0.0-rc.17
Install with support for both:
pip install simple_graph_retriever[all]==1.0.0-rc.17
2. Configuration
The SDK is configured via environment variables, which are loaded from a .env file in your project's root directory. You must configure either Neo4j or FalkorDB.
Common Configuration
| Variable | Default Value | Description |
|---|---|---|
GRAPH_DATABASE_TYPE |
neo4j |
The type of database to use: neo4j or falkordb. |
QDRANT_URL |
http://localhost:6333 |
The URL for your Qdrant vector database instance. |
QDRANT_API_KEY |
None |
Optional: The API key for authenticating with your Qdrant instance. |
QDRANT_CHUNKS_COLLECTION |
chunks |
The name of the collection for storing graph chunks. |
QDRANT_COMMUNITIES_COLLECTION |
communities |
The name of the collection for storing community summaries. |
EMBEDDER_URL |
http://localhost:8080 |
The URL of the text embedding service. It must accept a POST request with {"inputs": ["text"]}. |
VECTOR_SIZE |
384 |
The dimension of the vectors produced by your embedding model. |
LOGLEVEL |
INFO |
The logging level for the SDK (DEBUG, INFO, WARNING, ERROR). |
Neo4j Specific Configuration
| Variable | Default Value | Description |
|---|---|---|
NEO4J_URI |
bolt://localhost:7687 |
The URI for your Neo4j database. |
NEO4J_USER |
neo4j |
The username for your Neo4j database. |
NEO4J_PASSWORD |
password |
The password for your Neo4j database. |
FalkorDB Specific Configuration
| Variable | Default Value | Description |
|---|---|---|
FALKORDB_HOST |
localhost |
The host for your FalkorDB instance. |
FALKORDB_PORT |
6379 |
The port for your FalkorDB instance. |
FALKORDB_PASSWORD |
None |
Optional: Password for Redis/FalkorDB. |
FALKORDB_KEY |
graph |
The key (graph name) used in FalkorDB. |
Example .env files
Option A: Using Neo4j
GRAPH_DATABASE_TYPE="neo4j"
NEO4J_URI="bolt://localhost:7687"
NEO4J_USER="neo4j"
NEO4J_PASSWORD="your_secure_password"
QDRANT_URL="http://localhost:6333"
EMBEDDER_URL="http://localhost:8080/embed"
Option B: Using FalkorDB
GRAPH_DATABASE_TYPE="falkordb"
FALKORDB_HOST="localhost"
FALKORDB_PORT="6379"
QDRANT_URL="http://localhost:6333"
EMBEDDER_URL="http://localhost:8080/embed"
3. Usage
Initializing the Client
The main entry point is the GraphRetrievalClient. It automatically loads settings from your .env file and initializes the appropriate database driver (Neo4j or FalkorDB) based on your configuration.
from graph_retrieval_sdk.client import GraphRetrievalClient
# Initialize the client (backend determined by env vars)
client = GraphRetrievalClient()
Indexing the Graph
The index() method runs the full, idempotent pipeline to populate Qdrant with data from your connected graph database.
# This runs the full pipeline:
# 1. Community detection
# 2. Chunk creation
# 3. Chunk and Community indexing
client.index()
print("Indexing complete!")
To completely reset your vector index, use the clear_index() method. Warning: This is a destructive operation that deletes and recreates the Qdrant collections.
client.clear_index()
Retrieving a Subgraph
Use the retrieve_graph() method to query your indexed graph. The RetrievalConfig model allows for fine-grained control over the process.
Basic Retrieval
from graph_retrieval_sdk.models import RetrievalConfig
import json
query = "Who was the emperor of Rome?"
config = RetrievalConfig()
subgraph = client.retriever.retrieve_graph(query, config)
if subgraph:
print(f"Retrieved {len(subgraph[0]['nodes'])} nodes and {len(subgraph[0]['relationships'])} relationships.")
Advanced Retrieval with Graph Expansion Control
Customize the RetrievalConfig to tune the retrieval and expansion behavior.
from graph_retrieval_sdk.models import RetrievalConfig
query = "Tell me about the Roman emperors and their families, but exclude servants."
advanced_config = RetrievalConfig(
# --- Retrieval Settings ---
max_communities=5,
max_chunks=15,
# --- Graph Expansion Settings ---
max_hops=2,
community_expansion_limit=5,
# --- Relationship Filtering ---
allowed_rel_types=["HAS_SON", "MARRIED_TO", "SUCCESSOR_OF"],
denied_rel_types=["HAS_SERVANT"]
)
subgraph = client.retriever.retrieve_graph(query, advanced_config)
if subgraph:
print(f"Retrieved {len(subgraph[0]['nodes'])} nodes with advanced configuration.")
# Close the client connection when done
client.close()
4. Project Structure
The project is organized as a standard Python package:
├── graph_retrieval_sdk/
│ ├── __init__.py
│ ├── client.py # Main client class
│ ├── config.py # Configuration and logging setup
│ ├── indexer.py # Logic for indexing the graph
│ ├── retriever.py # Logic for retrieving subgraphs
│ ├── db/ # Database adapters
│ │ ├── neo4j.py
│ │ └── falkordb.py
│ └── models.py # Pydantic models (e.g., RetrievalConfig)
├── examples/
├── pyproject.toml # Project metadata and dependencies
└── README.md # This file
5. Contributing
Contributions are welcome! Please feel free to submit a pull request or open an issue.
6. License
This project is licensed under the MIT License.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file simple_graph_retriever-1.0.0rc17.tar.gz.
File metadata
- Download URL: simple_graph_retriever-1.0.0rc17.tar.gz
- Upload date:
- Size: 13.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
18eb12fd748d1bb77a81032a1ece71cca7190a264c3e68f65248d5fa6bc06f13
|
|
| MD5 |
c32cac525fb0de57f0405d4bf28769dc
|
|
| BLAKE2b-256 |
283b7ea5be1014bbd406fdbbddedc84df6ae5fda44856d24301b47dda3fe97a0
|
File details
Details for the file simple_graph_retriever-1.0.0rc17-py3-none-any.whl.
File metadata
- Download URL: simple_graph_retriever-1.0.0rc17-py3-none-any.whl
- Upload date:
- Size: 15.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5de5681515b05efa614d524e4419bbf8233a4f029cca70e93f3e5fa96a56fd2c
|
|
| MD5 |
03ae44506f054b64ed3a1ef63b03e07a
|
|
| BLAKE2b-256 |
6cc21d38e1931d3baf85098fa1dc61637d2c7420caba1ce61253870750f8a9e6
|