Skip to main content

Graph-based memory system using DuckDB

Project description

GraphMemory - GraphRAG Database

GraphMemory

Overview

An embedded graph database with vector similarity search (VSS), full-text search (BM25), and hybrid search using DuckDB. The GraphMemory class provides a complete API for managing nodes and edges, with a fluent query builder, graph import/export, connection pooling, and automatic retry logic.

Each node has a unique ID, a JSON properties field (any arbitrary dictionary), a node type (ex: Person, Organization, etc.), and a vector of floating point values.

Each edge has a unique ID, a source node ID, a target node ID, a relationship type (ex: served_under, worked_with, etc.), and a weight.

This database can be used for any graph-based RAG application or knowledge graph application.

Vector embeddings can be created using sentence-transformers or other API based models.

Features

  • Vector Similarity Search — HNSW-indexed nearest neighbor search with L2, cosine, and inner product distance metrics
  • Full-Text Search — BM25-scored text search across node properties
  • Hybrid Search — Combined vector + text search with configurable weights
  • DSPy Extraction — Automatic entity and relationship extraction from unstructured text using DSPy typed predictors
  • Graph Algorithms — PageRank, betweenness centrality, degree distribution, and connected components via NetworkX
  • Query Builder — Fluent, composable, parameterized query API with traversal support
  • Graph Import/Export — JSON, CSV, and GraphML formats
  • Connection Pooling — Thread-safe operations with automatic retry on transient errors
  • Context Manager — Use with statements for automatic resource cleanup

Installation

From PyPI

pip install graphmemory

Optional Dependencies

pip install graphmemory[extraction]   # DSPy-based entity/relationship extraction
pip install graphmemory[algorithms]   # Graph algorithms via NetworkX

From Source (Development)

git clone https://github.com/bradAGI/GraphMemory.git
cd GraphMemory
pip install -e .

Usage

Initialization

from graphmemory import GraphMemory, Node, Edge

# In-memory database
graph_db = GraphMemory(vector_length=1536)

# Persistent database with cosine distance
graph_db = GraphMemory(
    database='graph.db',
    vector_length=1536,
    distance_metric='cosine'  # 'l2', 'cosine', or 'inner_product'
)

# Using context manager
with GraphMemory(database='graph.db', vector_length=1536) as graph_db:
    # ... operations ...
    pass  # connection closed automatically

Auto Generated UUID

IDs for nodes and edges are auto generated UUIDs.

Query Builder

The GraphMemory class provides a fluent query builder via the query() method. All conditions use parameterized queries for safety.

# Find all Person nodes named George Washington
results = graph.query().match(type="Person").where(name="George Washington").execute()

# Traverse 2 hops from a node
results = graph.query().traverse(source_id=node_id, depth=2).execute()

# Paginate and order results
results = graph.query().match(type="Person").order_by("name").limit(10).offset(0).execute()

# Query edges from Person nodes
edges = graph.query().match(type="Person").edges().execute()

DSPy Extraction

The graphmemory.extraction module uses DSPy typed predictors to automatically extract entities (nodes) and relationships (edges) from unstructured text. Requires the extraction optional dependency.

from graphmemory import GraphMemory
from graphmemory.extraction import extract, extract_and_store
import dspy

# Configure your DSPy language model
dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))

text = """George Washington was the first President of the United States.
Thomas Jefferson served as Secretary of State under Washington.
Alexander Hamilton was the first Secretary of the Treasury."""

# Extract nodes and edges without storing
nodes, edges = extract(text)
for node in nodes:
    print(f"{node.type}: {node.properties}")

# Or extract and insert directly into a GraphMemory instance
graph_db = GraphMemory(vector_length=3)
inserted_nodes, edges = extract_and_store(graph_db, text)

Available functions:

Function Description
extract_nodes(text, sentences=None) -> list[Node] Extract entity nodes (proper nouns) from text.
extract_edges(text, nodes, sentences=None) -> list[Edge] Extract relationships between known nodes.
extract(text, sentences=None) -> tuple[list[Node], list[Edge]] Extract both nodes and edges in one call.
extract_and_store(graph, text, sentences=None) -> tuple[list[Node], list[Edge]] Extract and insert into a GraphMemory instance.

Graph Algorithms

The graphmemory.algorithms module provides graph analysis via NetworkX. Requires the algorithms optional dependency.

from graphmemory import GraphMemory
from graphmemory.algorithms import pagerank, betweenness_centrality, connected_components, to_networkx

graph_db = GraphMemory(vector_length=3)
# ... insert nodes and edges ...

# Compute PageRank scores
scores = pagerank(graph_db)

# Find betweenness centrality
centrality = betweenness_centrality(graph_db)

# Get connected components
components = connected_components(graph_db)

# Export to NetworkX for custom analysis
G = to_networkx(graph_db)

Available functions:

Function Description
to_networkx(graph) -> nx.DiGraph Export a GraphMemory instance to a NetworkX DiGraph.
pagerank(graph, alpha=0.85, max_iter=100, tol=1e-6) -> dict[str, float] Compute PageRank for every node.
betweenness_centrality(graph, normalized=True, endpoints=False) -> dict[str, float] Compute betweenness centrality for every node.
degree_distribution(graph) -> dict[str, dict[str, int]] Compute in/out/total degree for every node.
connected_components(graph) -> list[set[str]] Find weakly connected components (sorted largest-first).

Example Usage

from graphmemory import GraphMemory, Node, Edge

import json
from openai import OpenAI
import os

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

# Sample unstructured text
gw_text = "George Washington was the first President of the United States and served from 1789 to 1797."
tj_text = "Thomas Jefferson was the first Secretary of State of the United States and served from 1790 to 1793."
ah_text = "Alexander Hamilton was the first Secretary of the Treasury of the United States and served from 1789 to 1795."

# Extract structured data from unstructured text
def extract_attributes(text):
    return json.loads(client.chat.completions.create(
        model="gpt-3.5-turbo",
        messages=[
            {"role": "system", "content": "Extract structured data from this text using the following attributes: \
             name, title, country, term_start, term_end"},
            {"role": "user", "content": text}
        ],
        seed=1
    ).choices[0].message.content)

# Calculate embedding for a given input
def calculate_embedding(input_json):
    return client.embeddings.create(
        input=input_json,
        model="text-embedding-3-small"
    ).data[0].embedding

gw_embedding = calculate_embedding(gw_text)
tj_embedding = calculate_embedding(tj_text)
ah_embedding = calculate_embedding(ah_text)

# Initialize the database from disk (make sure to set vector_length correctly)
graph_db = GraphMemory(database='graph.db', vector_length=len(gw_embedding))

# Extract structured data from unstructured text
gw_attributes = extract_attributes(gw_text)
tj_attributes = extract_attributes(tj_text)
ah_attributes = extract_attributes(ah_text)

print(gw_attributes)
print(tj_attributes)
print(ah_attributes)

# Output Example:
# {
#   'person': 'George Washington',
#   'title': 'President',
#   'country': 'United States',
#   'term_start': '1789',
#   'term_end': '1797'
# }
# {
#   'person': 'Thomas Jefferson',
#   'title': 'Secretary of State',
#   'country': 'United States',
#   'term_start': 1790,
#   'term_end': 1793
# }
# {
#   'person': 'Alexander Hamilton',
#   'title': 'Secretary of the Treasury',
#   'country': 'United States',
#   'term_start': 1789,
#   'term_end': 1795
# }


# Create nodes with UUIDs
gw_node = Node(properties=gw_attributes, vector=gw_embedding)
tj_node = Node(properties=tj_attributes, vector=tj_embedding)
ah_node = Node(properties=ah_attributes, vector=ah_embedding)

gw_node_id = graph_db.insert_node(gw_node)
if gw_node_id is None:
    raise ValueError("Failed to insert George Washington node")

tj_node_id = graph_db.insert_node(tj_node)
if tj_node_id is None:
    raise ValueError("Failed to insert Thomas Jefferson node")

ah_node_id = graph_db.insert_node(ah_node)
if ah_node_id is None:
    raise ValueError("Failed to insert Alexander Hamilton node")

# Insert edges
edge1 = Edge(source_id=gw_node_id, target_id=tj_node_id, relation="served_under", weight=0.5)
edge2 = Edge(source_id=gw_node_id, target_id=ah_node_id, relation="served_under", weight=0.5)
graph_db.insert_edge(edge1)
graph_db.insert_edge(edge2)

# Print edges
print(graph_db.edges_to_json())

# Find connected nodes
connected_nodes = graph_db.connected_nodes(gw_node_id)
for node in connected_nodes:
    print("Connected Node Data:", node.properties)

# Find nearest nodes by vector embedding
nearest_nodes = graph_db.nearest_nodes(calculate_embedding("George Washington"), limit=1)
print(nearest_nodes)
print("Nearest Node Data:", nearest_nodes[0].node.properties)
print("Nearest Node Distance:", nearest_nodes[0].distance)

# Full-text search across node properties
results = graph_db.search_nodes("President", limit=5)
for result in results:
    print(f"Found: {result.node.properties} (score: {result.score})")

# Hybrid search (combines vector similarity + text search)
results = graph_db.hybrid_search(
    query_text="President",
    query_vector=calculate_embedding("President of the United States"),
    limit=5,
    text_weight=0.5,
    vector_weight=0.5
)

# Get node/s by attribute (Who was the Secretary of State?)
nodes = graph_db.nodes_by_attribute("title", "Secretary of State")
if nodes:
    print("Node by attribute:", nodes[0].properties)
else:
    print("No nodes found with the attribute 'title' = 'Secretary of State'")

# What is the title of the people who served under George Washington?
for node in connected_nodes:
    print(f"{node.properties.get('name')} - {node.properties.get('title')}")

# Fetch a node by UUID
fetched_node = graph_db.get_node(gw_node_id)

# Update a node
graph_db.update_node(gw_node_id, properties={"name": "George Washington", "title": "President"})

# Delete an edge by source / target node id
graph_db.delete_edge(edge1.source_id, edge1.target_id)

# Export the graph
graph_json = graph_db.export_graph(format='json')
graph_csv = graph_db.export_graph(format='csv')
graph_graphml = graph_db.export_graph(format='graphml')

Data Models

Node

Node(
    id: uuid.UUID,          # Auto-generated
    properties: dict | None, # Arbitrary key-value pairs
    type: str | None,        # Node type (e.g., "Person", "Organization")
    vector: list[float] | None  # Embedding vector
)

Edge

Edge(
    id: uuid.UUID,           # Auto-generated
    source_id: uuid.UUID,    # Source node ID
    target_id: uuid.UUID,    # Target node ID
    relation: str | None,    # Relationship type (e.g., "served_under")
    weight: float | None     # Edge weight
)

NearestNode

NearestNode(
    node: Node,       # The matched node
    distance: float   # Distance from query vector
)

SearchResult

SearchResult(
    node: Node,    # The matched node
    score: float   # Relevance score
)

TraversalResult

TraversalResult(
    node: Node,              # The reached node
    depth: int,              # Hops from source
    path: list[uuid.UUID]    # Node IDs in the path from source
)

GraphMemory API Reference

Initialization & Connection

Method Description
__init__(database=None, vector_length=3, distance_metric='l2', max_retries=3, retry_base_delay=0.1) Initialize the database. database: file path or None for in-memory. distance_metric: 'l2', 'cosine', or 'inner_product'.
close() Close the database connection (thread-safe).
cursor() Return a new DuckDB cursor for individual operations.
transaction() Context manager for database transactions.

Node Operations

Method Description
insert_node(node: Node) -> uuid.UUID Insert a node and return its ID.
bulk_insert_nodes(nodes: list[Node]) -> list[Node] Bulk insert multiple nodes.
get_node(node_id: uuid.UUID) -> Node Retrieve a node by ID.
update_node(node_id: uuid.UUID, **kwargs) -> bool Update node fields (type, properties, vector).
delete_node(node_id: uuid.UUID) Delete a node and its associated edges.
bulk_delete_nodes(node_ids: list[uuid.UUID]) Bulk delete multiple nodes and their edges.
nodes_by_attribute(attribute, value, limit=None, offset=None) -> list[Node] Query nodes by a property key-value pair.
get_nodes_vector(node_id: uuid.UUID) -> list[float] Retrieve the vector of a node.
nodes_to_json(limit=None, offset=None) -> list[dict] Export all nodes as JSON.

Edge Operations

Method Description
insert_edge(edge: Edge) Insert an edge between two nodes.
bulk_insert_edges(edges: list[Edge]) Bulk insert multiple edges.
get_edge(edge_id: uuid.UUID) -> Edge | None Retrieve an edge by ID.
get_edges_by_relation(relation: str) -> list[Edge] Get all edges with a given relation type.
edges_by_attribute(attribute: str, value) -> list[Edge] Query edges by attribute.
update_edge(edge_id: uuid.UUID, **kwargs) -> bool Update edge fields (relation, weight).
delete_edge(source_id: uuid.UUID, target_id: uuid.UUID) Delete an edge by source and target node IDs.
bulk_delete_edges(edge_ids: list[uuid.UUID]) Bulk delete multiple edges.
edges_to_json(limit=None, offset=None) -> list[dict] Export all edges as JSON.

Search & Similarity

Method Description
nearest_nodes(vector: list[float], limit: int) -> list[NearestNode] Find nearest neighbors by vector similarity.
search_nodes(query_text: str, limit: int = 10) -> list[SearchResult] Full-text search across node properties (BM25).
hybrid_search(query_text, query_vector, limit=10, text_weight=0.5, vector_weight=0.5) -> list[SearchResult] Combined text + vector search with configurable weights.
create_index() Create an HNSW index on node vectors for faster search.
set_vector_length(vector_length) Set the vector dimension for the database.

Graph Traversal

Method Description
connected_nodes(node_id: uuid.UUID) -> list[Node] Retrieve all nodes connected to a given node.
query() -> QueryBuilder Return a fluent query builder for composable, parameterized queries.

Import / Export

Method Description
export_graph(format='json') Export the graph. Formats: 'json', 'json_string', 'csv', 'graphml'.
import_graph(data, format='json') Import a graph from the given data and format.

Utility

Method Description
print_json() Print a JSON representation of all nodes and edges.

Examples

See the examples/ directory for complete usage examples:

  • openai_example.py — Uses OpenAI embeddings for node creation, similarity search, and attribute queries
  • lexical_graph.py — Wikipedia text extraction with SentenceTransformer embeddings
  • dspy_example_typed_pred.py — Knowledge graph extraction from unstructured text using DSPy

Testing

Unit tests are provided in tests/tests.py.

Running Tests

python3 -m unittest discover -s tests

License

This project is licensed under the MIT License. See the LICENSE file for details.

Contributing

Contributions are welcome! Please open an issue or submit a pull request.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

graphmemory-1.1.0.tar.gz (79.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

graphmemory-1.1.0-py3-none-any.whl (26.6 kB view details)

Uploaded Python 3

File details

Details for the file graphmemory-1.1.0.tar.gz.

File metadata

  • Download URL: graphmemory-1.1.0.tar.gz
  • Upload date:
  • Size: 79.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for graphmemory-1.1.0.tar.gz
Algorithm Hash digest
SHA256 4a98ca9616e6e042e9784059611dd9f92381e9025de5ace82e6c773b3473c23c
MD5 38cd5d8b6b0c557a82ffa66f71c2a511
BLAKE2b-256 b897bf5d53801033a5c67e0e17f91af5b552e4c917a506eba2a14466b157eb98

See more details on using hashes here.

File details

Details for the file graphmemory-1.1.0-py3-none-any.whl.

File metadata

  • Download URL: graphmemory-1.1.0-py3-none-any.whl
  • Upload date:
  • Size: 26.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for graphmemory-1.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 60ea2c5087d3888bd536d591044c47ab8988b8ceaca23454c8b073b36fbbcab0
MD5 e6701eb9092b0015a1b9de2dabc71e42
BLAKE2b-256 b4505554ea267c2f572e15ac3198b130356ac87680ec7130980b8b52c7158298

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page