treerag 🌳
RAG without vector databases.
A lightweight structural RAG library that indexes documents as hierarchical trees and retrieves answers using structure instead of embeddings.
Works with any LangChain-compatible LLM: OpenAI, Anthropic, Gemini, Ollama.
🚀 Why treerag?
- ❌ No vector database required
- ⚡ Fast hierarchical retrieval
- 🧠 Uses document structure instead of embeddings
- 🔌 Works with any LLM
- 🪶 Lightweight and easy to integrate
📦 Install
# Base
pip install treerag
# With OpenAI
pip install "treerag[openai]"
# With Anthropic
pip install "treerag[anthropic]"
# Everything
pip install "treerag[all]"
⚡ Quick Start
from treerag import index_document, make_summarizer, ask, make_retriever
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="gpt-4o")
# 1. Index document
doc = index_document("my_doc.md", summarizer=make_summarizer(llm))
# 2. Ask question
result = ask("What does this document cover?", doc, make_retriever(llm))
print(result.content) # answer
print(result.references) # sections used
print(result.response_metadata) # token usage
🧠 How It Works
Indexing
File / URL
↓ read_file()
↓ parse_sections()
↓ make_summarizer()
↓ build_hierarchy()
↓ flatten_tree()
↓ save_registry()
Querying
Query
↓ tree search (LLM selects relevant nodes)
↓ fetch content
↓ answer generation
↓ AIMessage with references
📂 Supported Inputs
# Files
index_document("file.md")
index_document("report.pdf")
index_document("document.docx")
index_document("notes.txt")
# URLs
index_document("https://docs.example.com")
📊 Response Formats
# Default (LangChain AIMessage)
result = ask("What is this?", doc, retriever)
print(result.content)
print(result.references)
# Plain dict
result = ask("What is this?", doc, retriever, return_raw=False)
print(result["answer"])
# Streaming
for chunk in ask("What is this?", doc, retriever, stream=True):
if isinstance(chunk, dict):
print(chunk["__references__"])
else:
print(chunk, end="")
⚡ Async Support
from treerag import aask, make_async_retriever
retriever = make_async_retriever(llm)
result = await aask("What is this?", doc, retriever)
print(result.content)
📚 Multi-Document Q&A
from treerag import ask_multi, get_document_by_id
doc1 = get_document_by_id("uuid-1")
doc2 = get_document_by_id("uuid-2")
result = ask_multi("What is the budget cap?", [doc1, doc2], retriever)
print(result.content)
print(result.references)
🗂️ Registry Management
from treerag import list_documents, get_document_by_id, delete_document
for doc in list_documents():
print(doc["name"], doc["doc_id"])
doc = get_document_by_id("uuid")
delete_document("uuid")
🎯 Custom Prompts
summarizer = make_summarizer(
llm,
system_prompt="You are a legal expert. Summarize clauses and obligations."
)
retriever = make_retriever(
llm,
answer_system_prompt=(
"You are a legal assistant. "
"Answer using ONLY the provided context."
)
)
result = ask(
"What are the key obligations?",
doc,
retriever,
extra_context="This is a legal agreement."
)
🔌 Supported Providers
from langchain_openai import ChatOpenAI
from langchain_anthropic import ChatAnthropic
from langchain_google_genai import ChatGoogleGenerativeAI
from langchain_ollama import ChatOllama
# OpenAI
make_retriever(ChatOpenAI(model="gpt-4o"))
# Anthropic
make_retriever(ChatAnthropic(model="claude-haiku"))
# Gemini
make_retriever(ChatGoogleGenerativeAI(model="gemini-2.0-flash"))
# Ollama (local)
make_retriever(ChatOllama(model="llama3"))
🏗️ Production Usage
doc = index_document("file.md", persist=False)
# store externally
db.save(doc["doc_id"], doc)
# load later
doc = db.get("id")
result = ask("Your question?", doc, retriever)
📌 Example Use Cases
- 📄 Documentation Q&A
- 📚 Internal knowledge base
- 🤖 AI assistants without vector DB
- 🧾 Legal / contract analysis
⚔️ treerag vs Traditional RAG
| Feature | treerag | Vector RAG |
|---|---|---|
| Setup | Simple | Complex |
| DB required | ❌ | ✅ |
| Cost | Low | High |
| Retrieval | Structure-based | Embeddings |
📄 License
MIT
Metadata
Release files for treerag 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| treerag-0.1.5.tar.gz | 19.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| treerag-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 39.2 kB
Release files / treerag-0.1.5.tar.gz
| Download URL | treerag-0.1.5.tar.gz |
|---|---|
| Size | 19.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ab66ff1c6c5777babffd7da852352c0b236caebfafb7e210439157f9494c4ba5
|
|
BLAKE2b-256 checksum How to use checksums |
db24b940613b0f5146da0817ae62c133aaee050ae10ba7f6868c28ee107a06ef
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.3
|
Release files / treerag-0.1.5-py3-none-any.whl
| Download URL | treerag-0.1.5-py3-none-any.whl |
|---|---|
| Size | 19.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e30b02187e07efebdb3c67e085bb0a439461e66173701e00a22b88c3acfe5491
|
|
BLAKE2b-256 checksum How to use checksums |
f941530350a5ac05f8714ac921c7c93b23a15f64dfd6955ea18cf44980f8e4e2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.3
|