Skip to main content

A framework that replace traditional RAG pipelines. Ingest any number of documents in multiple workspaces (channels, departments, etc.), index it with BM25, and let the agent search, fetch, and reason over it, exactly like searching the web, but entirely on your machine. No vector store, no embedding needed.

Project description

Local Search Agent

Give your AI agent a search engine for your local files.


What is this?

Local Search Agent is a Python framework that gives your AI agent a search engine for your local files and lets it search, fetch, and reason over your local documents — the same way a researcher searches the web, but entirely on your machine.

Point it at a folder. Ask a question. The agent searches your documents, reads the relevant ones, and gives you an answer with citations — no cloud upload, no API calls to external search services, no embeddings, no vector stores.

"What was the AWS spend in Q3?"  →  agent searches index  →  fetches relevant docs  →  answers with sources

Why not RAG?

Traditional RAG (Retrieval-Augmented Generation) has a fundamental problem: it converts your documents into embeddings and stores them in a vector database. That means:

  • Stale indexes — embeddings go out of date silently. You never know if the agent is reading your latest documents or a six-month-old snapshot
  • Black-box retrieval — you can't see why a document was retrieved or not. Debugging poor answers is guesswork
  • Chunking anxiety — split too small and you lose context. Split too large and retrieval quality degrades. There's no right answer
  • Infrastructure overhead — a vector database is another service to run, maintain, and pay for
  • Semantic drift — embeddings are sensitive to how questions are phrased. A question about "cloud expenditure" may never match a document that says "AWS spend"

Local Search Agent takes a different approach: BM25 keyword search via Meilisearch, structured metadata, and a LangGraph agent loop with tools. The agent searches your document index the same way a developer searches Stack Overflow — with real queries, real results, and full transparency into what was retrieved and why.

The result is deterministic, auditable, and fast. You can see exactly what the agent fetched for every answer.


How it works

1. INGEST     Your documents → parsed, cleaned, chunked, indexed into Meilisearch
2. SERVE      FastAPI file server makes documents available to the agent via HTTP
3. SEARCH     LangGraph agent loop: search_local_index → fetch_local_url → reason
4. ANSWER     Agent returns an answer with inline source citations

Everything runs locally. Meilisearch downloads automatically on first use, no manual setup.


Screenshots

Desktop UI

Local Search Agent UI

CLI Interactive Mode

Local Search Agent CLI

Python API

Local Search Agent Python API


Video Demos


Install

pip install local-search-agent

Set your API key

# Google AI Studio (free tier — recommended) or paid from openai or anthropic
local-search config set-key --provider google --key YOUR_KEY

# Or use Ollama for a fully local, zero-cost setup (no key needed)
# Install from https://ollama.com 
# Download any model that support function calling and system instructions: 
`ollama pull gemma4:e2b` (7.2GB) or `ollama pull gemma4:e4b` (9.6GB)

Quick Start

Desktop UI

local-search ui

The desktop UI open:

  1. Create a workspace, name it, point it at a directory of files.
  2. Ingest (parse, clean, chunk).
  3. Get a free google api key from ai-studio.
  4. Set your api key at the top bar's right corner, or add a paid key for anthropic\openai . Note: For paid models or ollama, you will need to set model name via the config button at the top bar's right corner.
  5. click Ingest from the left sidebar.
  6. watch the progress bar at the bottom bar, wait until all files marked as completed.
  7. Start asking questions.

CLI

# Create a workspace and ingest documents
local-search workspace create finance "C:\my_docs"
local-search ingest --workspace finance --dirs "C:\my_docs"

# Start the file server (keep this running)
local-search serve --workspace finance

# Ask a question
local-search query "What was the AWS spend in Q3?" --workspace finance --provider google

# Use interactive mode
local-search --workspace finance --provider google

Python API

from local_search_agent import SearchAgentFramework, SearchAgentConfig

config = SearchAgentConfig(
    document_dirs=["C:/my_docs"],
    workspace_name="finance",
    provider="google",
)

framework = SearchAgentFramework(config)
framework.ingest_and_index()
framework.start_file_server()

response = framework.query("What was the AWS spend in Q3?")
print(response["answer"])

Supported File Types

Format Extension
PDF .pdf
Word .docx
Excel .xlsx
PowerPoint .pptx
HTML .html, .htm
Plain text .txt, .md
CSV .csv
JSON .json
XML .xml
Email .eml

Key Features

  • One command installpip install local-search-agent. Meilisearch downloads automatically
  • No embeddings, no vector stores — BM25 search with structured metadata. Fast, deterministic, auditable
  • Native desktop UI — pywebview window with live streaming agent responses, workspace management, and chat history
  • Multi-provider LLM — Google, Ollama (local), OpenAI, Anthropic
  • Multi-workspace — isolate document collections by department, project, channel, or topic. Each workspace is its own search index
  • Incremental sync — background scheduler re-indexes only changed files. A 10,000-document corpus with 50 changes re-indexes only the 50
  • Full CLI parity — everything you can do in the UI you can do from the terminal
  • Python API — embed the framework directly in your own application
  • Cross-platform — Windows, macOS, Linux

Documentation

Guide Description
Getting Started First steps, quick start for UI, CLI, and Python API
Installation Full install guide, API keys, Ollama setup, platform notes
Architecture Full architrecture, design guide
CLI Reference All commands and flags
Python API Reference Full API documentation
Configuration All config options and patterns
Ingestion How ingestion works, supported formats, chunking, scheduler
Multi-Workspace Managing multiple document collections
Semantic Search Experimental: concept extraction, query expansion, link graph
Troubleshooting Common issues and fixes

Contributing

Contributions are welcome. Clone the repo and install in editable mode with dev dependencies:

git clone https://github.com/wiss84/local-search-agent.git
cd local-search-agent
pip install -e ".[dev]"

Run tests before submitting a PR:

pytest tests/ -v --cov=local_search_agent --cov-report=term-missing
ruff check .
ruff format .

License

MIT — see LICENSE for details.


Built by Wissam Metawee

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

local_search_agent-0.1.2.tar.gz (1.8 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

local_search_agent-0.1.2-py3-none-any.whl (1.8 MB view details)

Uploaded Python 3

File details

Details for the file local_search_agent-0.1.2.tar.gz.

File metadata

  • Download URL: local_search_agent-0.1.2.tar.gz
  • Upload date:
  • Size: 1.8 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for local_search_agent-0.1.2.tar.gz
Algorithm Hash digest
SHA256 011e32111785c4b704bbbdb56e75ad50d1b4b42032d65c355c811311a5a9365e
MD5 22e1dea514999cc88bbdc5110b6e65e8
BLAKE2b-256 cca71bdfe74d2d50e5ad82e071b4828575de7a86fbf62c6711bdcce6e55182ee

See more details on using hashes here.

Provenance

The following attestation bundles were made for local_search_agent-0.1.2.tar.gz:

Publisher: ci-cd.yml on wiss84/local-search-agent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file local_search_agent-0.1.2-py3-none-any.whl.

File metadata

File hashes

Hashes for local_search_agent-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 65409275f4040cbc5b083ff5a4e4f3fcbaf41d4be9e45181fa52ffdc7c321d1b
MD5 577a39cfa3e92d3c6d380231fdb5ea35
BLAKE2b-256 0c7806309130452349aa7003222385ad385683c5338a81b280aa22ead7d2f860

See more details on using hashes here.

Provenance

The following attestation bundles were made for local_search_agent-0.1.2-py3-none-any.whl:

Publisher: ci-cd.yml on wiss84/local-search-agent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page