Skip to main content

🌌 Lokum Engine 🌟

The Undisputed King of RAG & LLM Fine-Tuning

PyPI Python License Status

From local experimentation to Fortune 500 production in 3 lines of code.

Quickstart • Features • Roadmap • Documentation


⚡ Why Lokum Engine?

Lokum Engine is the developer-first building block for Retrieval-Augmented Generation (RAG) and State-of-the-Art LLM Fine-Tuning. We abstracted away the infrastructure headaches, OOM crashes, and broken data pipelines so you can focus on building intelligent agents.

🚀 For Developers

  • Drop-in Simplicity: Setup RAG or start an MLX LoRA training loop in 3 lines of Python.
  • Quality Profiles: Sensible, pre-tuned defaults (base, mid, fab) that automatically balance speed vs. quality.
  • Fail-Fast Reliability: Strict data validation, deleted file reconciliation, and explicit error reporting. No silent failures.

🏢 For Enterprises

  • ChatML-Safe Presplitting: Guarantee your fine-tuning data never splits across critical instruction boundaries.
  • Persistent RAG State: Robust chunk tombstoning, metadata validation, and persistent storage.
  • Hardware Aware: Automatically detects and leverages Apple Silicon (MPS) and optimizes batch sizes.

📦 Install

pip install lokum-engine

(Note: Lokum Engine intentionally includes heavy, production-grade dependencies like FAISS, sentence-transformers, PyMuPDF, and MLX out of the box).


🧠 Quickstart: RAG (Retrieval-Augmented Generation)

Turn any folder of documents into a highly accurate semantic search engine instantly.

from lokum_engine import RAGEngineFab

# Initialize with the 'Fab' profile for maximum enterprise-grade retrieval quality
rag = RAGEngineFab()  

# Recursively ingest PDFs, Markdown, Code, and text files
rag.ingest_folder("/path/to/your/enterprise/docs", recursive=True)

# Query with semantic understanding
context = rag.query("How do we scale our distributed training pipeline?", k=5)
print(context)

🎯 Quickstart: Fine-Tuning (MLX LoRA)

Train state-of-the-art models on your own data without wrestling with CUDA errors or dataset corruption.

from lokum_engine import FinetuneEngineFab

# Initialize the engine
ft = FinetuneEngineFab(model_path="/path/to/mlx/base-model")

# Safely presplit the dataset to avoid OOMs while perfectly preserving ChatML tags
ft.presplit_dataset(
    dataset_dir="/path/to/raw/data", 
    max_seq_length=2048, 
    batch_size=4
)

# Launch the training loop
process = ft.start_training(
    dataset_path="/path/to/raw/data",
    batch_size=4,
    num_layers=16,
    iters=1000,
)

print(f"🚀 Training launched successfully! PID: {process.pid}")

🎛️ Quality Profiles: The Magic of Lokum

Stop guessing hyper-parameters. Lokum Engine ships with three tuned profiles for both RAG and Fine-Tuning:

Profile Target Audience Focus RAG Behavior Fine-Tune Behavior
Base Local Devs Speed & Efficiency Lighter embedding models, faster retrieval Smaller batch sizes, faster epochs
Mid Startups The Sweet Spot Balanced chunking and embedding Standard LoRA parameters
Fab Enterprises Maximum Quality Heavy embeddings, aggressive retrieval High-layer targeting, max context length

🗺️ The Master Roadmap

We are on a mission to become the industry standard. Here is a sneak peek at what's next:

  • Hybrid Search & Re-ranking: BM25 + Dense embeddings sorted by Cohere/BGE.
  • Enterprise Vector DBs: Native Milvus, Pinecone, and Qdrant support.
  • RAG & Fine-Tune Eval: Built-in LLM-as-a-judge to measure precision and recall.
  • DPO / PPO Support: Move beyond SFT and align models with human preferences natively.

👉 View the full Master Roadmap here


🤝 Contributing & Community

Lokum Engine is built by developers, for developers. We welcome PRs, issues, and ideas. If this project helped you build something awesome, please leave a ⭐ on GitHub!

📜 License

MIT License - free for indie hackers and Fortune 500s alike.

Metadata

Release files for lokum-engine 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lokum-engine 1.0.0
File Size Uploaded
lokum_engine-1.0.0.tar.gz 46.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lokum-engine 1.0.0
File Interpreter ABI Platform
lokum_engine-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 84.5 kB

Release files / lokum_engine-1.0.0.tar.gz

Download URL lokum_engine-1.0.0.tar.gz
Size 46.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a810255b7d8a1a85e1dea6befba4cd15ea76f85b516c3845e18a9a11593b420c
BLAKE2b-256 checksum
How to use checksums
6a89640405358cf18b43d90c4092daca77baa96088333b4353ddb895bfc9a703
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release files / lokum_engine-1.0.0-py3-none-any.whl

Download URL lokum_engine-1.0.0-py3-none-any.whl
Size 38.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c43d04fbee5692a36730613b5eaf2fc965086920e1c5cd2d04efc4ff66283767
BLAKE2b-256 checksum
How to use checksums
c7418dc9334167cc72c2da06ae513b1c2ac4a4abc91adaaf6c3970e771605535
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page