Lightweight RAG library for internal knowledge systems
Project description
Fullstack RAG AI
Lightweight Retrieval-Augmented Generation (RAG) library for building internal knowledge systems using local documents and LLMs.
Supports:
- Vector-based retrieval using FAISS and embeddings
- Vectorless retrieval using BM25
Optional LLM integration allows reranking, query expansion, and dynamic prompt responses.
Features
- Load documents from PDFs or raw text
- Automatic chunking (per-page or whole PDF)
- Vector-based retrieval with FAISS + embeddings
- Vectorless retrieval with BM25
- Optional LLM reranking and query expansion
- Customizable prompts for LLMs
- Deterministic caching to avoid repeated LLM calls
- Automatic detection of document updates, additions, and deletions
- Fully modular: swap retriever, reranker, or query expander
Supported tools:
- Ollama LLMs
- HuggingFace embeddings
- FAISS vector database
Installation
pip install fullstack-rag-ai
Default Configuration
The library has two separate default configurations depending on the retrieval method.
- Vector-Based (FAISS + Embeddings)
DEFAULT_CONFIG = {
"embedding_model": "sentence-transformers/all-MiniLM-L6-v2",
"llm_model": "llama3",
"chunk_size": 500,
"chunk_overlap": 100,
"k": 15,
"chunking_strategy": "recursive",
"prompt_template": """
You are a helpful assistant. Use the context below to answer the question.
Combine information, reason, and infer if needed.
If answer is not in context, say "I could not find the answer in the documents."
Context:
{context}
Question:
{question}
"""
}
- Vectorless BM25
DEFAULT_CONFIG_BM25 = {
"top_k": 5,
"max_context_length": 2000,
"k1": 1.5, # BM25 parameter
"b": 0.75, # BM25 parameter
"reranker_prompt": """
You are a ranking system.
Query:
{query}
Documents:
{documents}
Return ONLY a list of indices (most relevant first)
""",
"query_expander_prompt": """
Expand the following query into 3 alternative search queries.
Query: {query}
Return as a comma-separated list.
"""
}
- top_k: number of documents retrieved per query
- max_context_length: max context length for LLM
- k1 & b: BM25 scoring parameters
Core Workflow
Vector-Based (FAISS + Embeddings)
Documents → Chunking → Embeddings → FAISS → Retrieval → LLM → Answer
Vectorless (BM25)
Documents → Chunking → BM25 Index → Retrieval → Optional LLM Rerank → Context → LLM Answer
Vector-Based FAISS Pipeline
Vector Database Synchronization
from fullstack_rag_lib.vector_rag_ai import VectorDBSynchronizer, QAService
# Step 1: Build or update vector database
syncer = VectorDBSynchronizer(
documents_path="./data",
index_path="./vector_db",
embedding_model="sentence-transformers/all-MiniLM-L6-v2",
chunk_size=600,
chunk_overlap=120,
chunking_strategy="semantic"
)
all_docs, metadata, qa_cache, message = syncer.sync()
print(message)
# Step 2: Ask questions
qa = QAService(index_path="./vector_db", model="llama3", k=15, debug=True)
answer = qa.ask("What is FAISS used for?")
print(answer)
- Supports custom prompt templates:
custom_prompt = """
You are a technical assistant. Use the context to answer concisely.
If answer is not in context, say 'Not found'.
Context:
{context}
Question:
{question}
"""
qa = QAService(index_path="./vector_db", prompt_template=custom_prompt)
Vectorless BM25 Pipeline
from fullstack_rag_ai.vectorless_rag_ai import VectorlessRAGPipeline
from fullstack_rag_ai.vectorless_rag_ai import load_pdfs
from fullstack_rag_ai.vectorless_rag_ai import VectorlessConfig
from fullstack_rag_ai.vectorless_rag_ai import LLMReranker
from fullstack_rag_ai.vectorless_rag_ai import QueryExpander
# Load documents
documents = load_pdfs("./data", chunk_pages=True)
# Initialize pipeline
pipeline = VectorlessRAGPipeline(
config=VectorlessConfig(top_k=5, max_context_length=2000)
)
pipeline.add_documents(documents)
# Optional LLM reranker
def my_llm(prompt: str) -> str:
return "0,1,2" # Replace with actual LLM call
pipeline.set_reranker(LLMReranker(my_llm))
# Optional query expansion
pipeline.set_query_expander(QueryExpander(my_llm))
# Run query
result = pipeline.run("tell me about ec2 instance m7a.medium cost")
print(result["prompt"])
print([r.chunk.id for r in result["results"]])
- Works without embeddings or FAISS
- Fully modular: swap retriever, reranker, or query expander dynamically
- Cache results with pipeline.clear_cache()
Changing the Embedding Model
Embeddings:
syncer = VectorDBSynchronizer(
documents_path="./data",
index_path="./vector_db",
embedding_model="sentence-transformers/all-mpnet-base-v2"
)
all_docs, metadata, qa_cache, message = syncer.sync()
LLM:
qa = QAService(
index_path="./vector_db",
model="llama3:8b",
embedding_model="sentence-transformers/all-MiniLM-L6-v2"
)
Any Ollama or locally available LLM can be used To install a model:
ollama pull llama3
ollama pull llama3:8b
ollama pull llama3:70b
ollama pull phi3
ollama pull mistral
Rebuilding Vector Database
If chunking or embedding model changes, rebuild in a new path:
syncer = VectorDBSynchronizer(
documents_path="./data",
index_path="./vector_new_db",
embedding_model="BAAI/bge-base-en",
chunk_strategy="token",
chunk_size=600,
chunk_overlap=120
)
all_docs, metadata, qa_cache, message = syncer.sync()
Deterministic Caching
Cache keys: question + retrieved context hash If question and context are unchanged, cached answer is returned Stored in qa_cache.bin
Dependencies
- Python >= 3.9
- pypdf (PDF loading)
- faiss-cpu (vector-based retrieval, optional if using BM25)
- sentence-transformers (embedding models, optional if using BM25)
- langchain and related packages for LLM integration:
- langchain
- langchain-community
- langchain-huggingface
- langchain-ollama
- langchain-experimental (optional)
Supports dynamic usage of:
- Any HuggingFace embedding model
- Any Ollama LLM model
License
MIT License
Author
Fullstack-Solutions
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file fullstack_rag_ai-1.0.0.tar.gz.
File metadata
- Download URL: fullstack_rag_ai-1.0.0.tar.gz
- Upload date:
- Size: 16.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
26b165f4ede77898537d31b5ca2a45352b2319dabe67b72d4b95d639a2136664
|
|
| MD5 |
50aa4d40c9ecfe10ab3ab02964d8644d
|
|
| BLAKE2b-256 |
aadad614539ceae1dd684b00b792d567fe66fecce534252b21b9ea3a12231d5a
|
File details
Details for the file fullstack_rag_ai-1.0.0-py3-none-any.whl.
File metadata
- Download URL: fullstack_rag_ai-1.0.0-py3-none-any.whl
- Upload date:
- Size: 21.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7445f749f8ac7bd3f482690a8bd9aa501d0908131a995f938d1e8bc18e10e10e
|
|
| MD5 |
fdc5a8a982688723e35815a02140b419
|
|
| BLAKE2b-256 |
72e577ab06a5fd6b13539606185a7c72576052d140976283c655323525bac16c
|