Heta
Heta is a Python framework for building, querying, and evaluating knowledge bases from composable recipes.
Why Heta
Most RAG projects become hard to change when parsing, models, storage, indexing, retrieval, and evaluation logic are wired directly into application code. Heta keeps those parts explicit and replaceable:
- Models wrap LLM and embedding providers through LiteLLM.
- Stores give objects, vectors, SQL rows, and graph facts consistent APIs.
- Steps describe one build action and the search capability it unlocks.
- Recipes compose models, stores, parsers, and steps into a reusable build plan.
- KnowledgeBase runs a recipe and exposes the search modes that were built.
- Benchmarks evaluate a recipe by building KBs, running queries, and writing reports.
Heta is not a fixed RAG pipeline. A recipe is the unit of composition: choose the components, choose the steps, build the KB, and run the same recipe against real benchmarks.
Install
pip install heta-framework
Optional extras:
pip install "heta-framework[sql]" # SQLStore and sql_text_search with SQLite or generic SQL
pip install "heta-framework[postgres]" # PostgreSQL driver for SQLStore and PostgreSQL text ranking
pip install "heta-framework[mysql]" # MySQL driver for SQLStore
pip install "heta-framework[milvus]" # Milvus vector store
pip install "heta-framework[s3]" # S3-compatible object store
pip install "heta-framework[text-index]" # Elasticsearch-backed full_text_search
Set a model key. Heta uses LiteLLM model names:
export OPENAI_API_KEY="sk-..."
Quick Start
This example builds a small vector-search KB from plain text, then asks the query engine to synthesize an answer from the retrieved evidence.
import asyncio
import os
from pathlib import Path
from heta_framework.common.models import EmbeddingModel, LanguageModel
from heta_framework.common.stores import InMemoryVectorStore, LocalObjectStore
from heta_framework.kb import (
DocumentParserRegistry,
EmbedChunks,
IndexVectors,
KnowledgeBase,
KnowledgeModels,
KnowledgeParsers,
KnowledgeRecipe,
KnowledgeStores,
ParseDocuments,
SplitDocuments,
TextParser,
)
async def main() -> None:
workspace = Path("heta-workspace")
objects = LocalObjectStore(workspace / "objects")
vectors = InMemoryVectorStore()
await objects.put(
"raw/aerospace-notes.txt",
(
"Heta builds knowledge bases from recipes. "
"A recipe combines parsers, models, stores, and steps. "
"Vector search retrieves chunks by semantic similarity."
).encode("utf-8"),
)
llm = LanguageModel(
model_name="openai/gpt-4o-mini",
api_key=os.environ["OPENAI_API_KEY"],
)
embedding = EmbeddingModel(
model_name="openai/text-embedding-3-small",
api_key=os.environ["OPENAI_API_KEY"],
)
recipe = KnowledgeRecipe(
parsers=KnowledgeParsers(
documents=DocumentParserRegistry([TextParser()]),
),
models=KnowledgeModels(language=llm, embedding=embedding),
stores=KnowledgeStores(objects=objects, vector=vectors),
steps=(
ParseDocuments(),
SplitDocuments(),
EmbedChunks(),
IndexVectors(),
),
)
kb = await KnowledgeBase.create(recipe=recipe, name="aerospace-notes")
response = await kb.query(
"How does Heta build a knowledge base?",
mode="vector_search",
options={"generate_answer": True},
)
print(kb.run_record.status)
print(sorted(kb.available_queries))
print(response.answer)
print(response.results[0].text)
await llm.aclose()
await embedding.aclose()
await vectors.aclose()
await objects.aclose()
asyncio.run(main())
Core Concepts
| Concept | Role |
|---|---|
KnowledgeRecipe |
Declares components and ordered build steps. |
KnowledgeBase |
Created from a recipe; exposes available query modes. |
ParseDocuments |
Converts raw objects into Heta ParsedDocument records. |
SplitDocuments |
Converts parsed documents into stable chunks. |
EmbedChunks |
Creates embeddings for chunks. |
IndexVectors |
Writes chunk vectors and unlocks vector_search. |
PersistChunks |
Writes chunk text to SQL and unlocks sql_text_search. |
IndexFullText |
Writes chunk text to a full-text index and unlocks full_text_search. |
HetaGraphProcedure |
Expands into entity, relation, and graph build steps. |
Build Patterns
The examples below are small recipes, not presets. They show how adding or
removing steps changes what the resulting KnowledgeBase can do. Any valid
recipe can be built, queried, deleted, and evaluated through the same framework
interfaces.
1. Vector Search
Use this for the smallest semantic-search KB. The recipe only needs an object store, an embedding model, a vector store, and the indexing steps.
recipe = KnowledgeRecipe(
parsers=KnowledgeParsers(documents=DocumentParserRegistry([TextParser()])),
models=KnowledgeModels(embedding=embedding),
stores=KnowledgeStores(objects=objects, vector=vectors),
steps=(
ParseDocuments(),
SplitDocuments(),
EmbedChunks(),
IndexVectors(),
),
)
kb = await KnowledgeBase.create(recipe=recipe, name="vector-kb")
response = await kb.query("knowledge base recipe", mode="vector_search")
Unlocked queries:
available queries: vector_search
2. Vector + SQL Text Search
Add SQLStore and PersistChunks when exact terms, product codes, legal
clauses, or operational phrases matter. This is the same recipe shape with one
extra store and one extra step.
from heta_framework.common.stores import SQLStore
from heta_framework.kb import PersistChunks, PersistChunksConfig
sql = SQLStore("sqlite:///heta-workspace/chunks.db")
recipe = KnowledgeRecipe(
parsers=KnowledgeParsers(documents=DocumentParserRegistry([TextParser()])),
models=KnowledgeModels(embedding=embedding),
stores=KnowledgeStores(objects=objects, vector=vectors, sql=sql),
steps=(
ParseDocuments(),
SplitDocuments(),
EmbedChunks(),
IndexVectors(),
PersistChunks(PersistChunksConfig(chunk_keys_artifact="chunk_keys")),
),
)
kb = await KnowledgeBase.create(recipe=recipe, name="keyword-kb")
semantic = await kb.query("safety checklist", mode="vector_search")
keyword = await kb.query("aerodynamic stall recovery", mode="sql_text_search")
Unlocked queries:
available queries: sql_text_search, vector_search
3. Heta Graph
Add a language model, SQL store, and HetaGraphProcedure when the KB should
extract entities and relations, then search graph facts with evidence. The graph
procedure is still just a group of steps, so teams can replace or extend it with
their own graph-building procedure.
from heta_framework.common.models import LanguageModel
from heta_framework.common.stores import SQLStore
from heta_framework.kb import HetaGraphProcedure
llm = LanguageModel(
model_name="openai/gpt-4o-mini",
api_key=os.environ["OPENAI_API_KEY"],
)
sql = SQLStore("sqlite:///heta-workspace/graph.db")
recipe = KnowledgeRecipe(
parsers=KnowledgeParsers(documents=DocumentParserRegistry([TextParser()])),
models=KnowledgeModels(language=llm, embedding=embedding),
stores=KnowledgeStores(objects=objects, vector=vectors, sql=sql),
steps=(
ParseDocuments(),
SplitDocuments(),
EmbedChunks(),
IndexVectors(),
*HetaGraphProcedure.build().steps(),
),
)
kb = await KnowledgeBase.create(recipe=recipe, name="graph-kb")
facts = await kb.query("What entities are connected?", mode="heta_graph_search")
Unlocked queries:
available queries: heta_graph_search, hybrid_search, vector_search
Evaluate Recipes
Benchmarks are built around recipes. A benchmark can create one KB for the whole corpus or many KBs for case-scoped documents, run the configured query modes, and write an evaluation report. This makes a recipe easy to compare before it is used in an application.
The benchmark runner has been exercised with real data:
| Benchmark | Scope | Verified path |
|---|---|---|
| UDA-fin | 788 PDFs, 8190 cases, multi-KB by doc_name |
Recipe build, source isolation, query, report output. |
| BEIR/SciFact | 5183 documents, 300 queries | Corpus-level KB build, vector search, metric report. |
| MultiHop-RAG | 609 documents, 2556 queries | Corpus-level KB build, vector search, evidence recall report. |
Swappable Components
The quick examples use local stores so they run anywhere. Production recipes can swap providers and storage backends without changing the step structure:
objects = S3ObjectStore(...)
vectors = MilvusVectorStore(...)
sql = SQLStore("postgresql+psycopg://user:password@host:5432/db")
The naming of tables, collections, and object prefixes stays explicit. Heta does not hide deployment boundaries from the recipe author.
Development
git clone https://github.com/KnowledgeXLab/Heta_Framework.git
cd Heta_Framework
pip install -e ".[dev]"
pytest
Build docs:
mkdocs serve
Metadata
Release files for heta-framework 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| heta_framework-0.1.1.tar.gz | 264.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| heta_framework-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 536.3 kB
Release files / heta_framework-0.1.1.tar.gz
| Download URL | heta_framework-0.1.1.tar.gz |
|---|---|
| Size | 264.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c04aa6b359314ed6ec95ffb31098a98021c7bf984c41e8eb4a088574ee93668e
|
|
BLAKE2b-256 checksum How to use checksums |
4b0a8c5be0f719f00d1f95318e4743577596021201ca31ac2e2e53743cf4b34d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.
Transparency logRelease files / heta_framework-0.1.1-py3-none-any.whl
| Download URL | heta_framework-0.1.1-py3-none-any.whl |
|---|---|
| Size | 271.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0b55af0aec492149e86bf0dd13605163f9d9176632972c76d3670ce463358475
|
|
BLAKE2b-256 checksum How to use checksums |
b852dab14468abb4c1604778e77e2c98cd1aa37bd25f1c337c554412122e3cde
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.
Transparency log