Skip to main content

AIAR logo

AIAR — Local RAG with LLM-as-judge and a grounding loop

AIAR is a local-first retrieval-augmented generation framework for Python. It runs against your own Ollama instance, ingests your own documents, and ships three production-grade primitives out of the box: hybrid retrieval, an LLM-as-judge that returns a structured Verdict, and a grounding store that lets accepted corrections feed back into future answers. It is built for developers, researchers, and AI hobbyists who want to own their stack end to end — no cloud calls, no telemetry, no vendor lock-in.

Install

pip install aiar-rag
# or, with the full retrieval extras (BM25, cross-encoder reranker, HyDE):
pip install 'aiar-rag[rag]'

Note on the name: the distribution on PyPI is aiar-rag because aiar was already taken. The import package remains aiar — so your code uses import aiar, but you install with pip install aiar-rag.

Prerequisites: Python 3.10+ and a running Ollama daemon (default http://127.0.0.1:11434). Pull at least one chat model and one embedding model, for example:

ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text

Quickstart

from aiar.harness.pipeline import answer_prompt

result = answer_prompt("What did our Q3 deployment doc say about rollback?", judge=True)
print(result["answer"])
print(result["verdict"])   # {"label": "Supported" | "Unsupported" | ..., "rationale": "...", ...}

result also carries grounded, reground_applied, retrieval, and latency_ms so you can wire the loop into your own UI or pipeline.

Why AIAR

Most local-RAG stacks stop at "retrieve and stuff into a prompt." AIAR treats the answer as the beginning of the loop, not the end. The judge catches hallucinations the moment they happen; the grounding store makes sure the same hallucination does not happen twice. The whole system runs on a laptop with a Qwen-class model and no external API calls — which means you can ship it into environments where cloud calls are not allowed, and you can audit every byte the model sees.

The three wedges

Hybrid retrieval

AIAR fuses lexical and semantic retrieval rather than picking one. Every query runs through BM25 over a tokenized index and a vector search over Ollama embeddings; the two ranked lists are merged with reciprocal-rank fusion (RRF), then optionally reranked by a cross-encoder. HyDE-style query rewriting and configurable top_k / fetch_k give you knobs without forcing a tuning project on day one.

LLM-as-judge Verdict

Every answer can be graded by a second LLM call that returns a structured Verdict: a label (Supported, Partially supported, Unsupported, Off-topic), a rationale, and the citations actually relied on. The judge sees the same retrieved context as the answerer, so its critique is grounded in evidence — and downstream code can branch on verdict.label to gate, retry, or escalate.

Grounding loop

When a Verdict is accepted (by a human, by automation, or by policy), the answer plus its supporting context is persisted to a grounding store keyed on the prompt. Next time a similar prompt arrives, AIAR reinjects that grounding before the answerer runs and flags reground_applied=True. The system stops re-making the same mistake — your corrections compound.

What is in the box

  • aiar.harness.pipeline.answer_prompt — the one-call entry point used in the Quickstart above.
  • aiar.rag — hybrid retrieval, BM25 + vector + RRF, optional cross-encoder reranker, HyDE rewriting.
  • aiar.eval — the LLM-as-judge with structured Verdict schema.
  • aiar.grounding — accepted-correction store and reground pipeline.
  • aiar.harness.service — optional FastAPI service exposing /services/prompt and /services/meta for other apps on the box.
  • aiar.observability — call IDs, latency, and retrieval traces for every answer.

Integration contracts (for apps built on AIAR)

AIAR exposes stable, schema-versioned contracts so a consuming app can use the framework's truth (retrieval, grounding, ingest readiness, telemetry) instead of reimplementing it. Each shape has one serializer shared by the in-process Python API and the HTTP route, so local and remote callers get identical payloads. See docs/integration-contracts.md for the frozen field-level shapes.

  • Pure retrieval (aiar.retrieve.v1) — aiar.rag.retrieve_chunks(query, *, instance, k=8, category=None) and GET/POST /instances/{instance}/retrieve. Raw vector similarity; never invokes a generation model.
  • Grounding (aiar.grounding.v1) — aiar.grounding.record_grounding(...) / lookup_grounding(...), keeping answer and correction distinct and scoping records per instance. Legacy records remain readable.
  • Ingest result + readiness (aiar.ingest.v1) — aiar.rag.ingest_documents( documents, *, instance, publish=False); health(instance) reports published, chunk_count, last_ingest_at, last_ingest_error.
  • Telemetry + capability manifest — answer_prompt(..., include_sources=True) attaches the answerer's source set; GET /capabilities returns the feature manifest a consumer gates on (never a version string); GET /calls/{call_id} returns a redacted trace; /healthz carries retrieve_schema_version.
  • Active-model reliability — active_model_ready on /healthz & /services/meta and features.generation on /capabilities flag when the configured model isn't pulled; generation then returns a structured model_not_pulled 4xx (not an opaque 503). Repoint live with authed POST /services/model, or opt into AIAR_ACTIVE_MODEL_FALLBACK=auto.

Configuration

AIAR reads AIAR_* environment variables for runtime configuration — endpoints, model names, reranker toggles, grounding-store paths, instance isolation, and so on. Sensible defaults work out of the box for a local Ollama install; see PLAYBOOK.md for the full matrix.

NEXT STEPS

Tune retrieval quality with the worked guides in examples/feature-guides/, starting with improving-rag.md.

Contributing

PRs welcome. The deep-dive operator guide lives at PLAYBOOK.md — end-to-end walkthrough covering ingestion, the harness, the watcher GUI, regrounding, evals, and operational notes. Worked examples live under examples/feature-guides/improving-rag.md.

For framework-level discussion, file an issue.

License

Apache-2.0. See LICENSE, NOTICE, or the license field in pyproject.toml.

Metadata

Release files for aiar-rag 0.2.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aiar-rag 0.2.5
File Size Uploaded
aiar_rag-0.2.5.tar.gz 725.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aiar-rag 0.2.5
File Interpreter ABI Platform
aiar_rag-0.2.5-py3-none-any.whl Python 3 none any Details

Total release size: 1.4 MB

Release files / aiar_rag-0.2.5.tar.gz

Download URL aiar_rag-0.2.5.tar.gz
Size 725.6 kB
Tags Source
SHA-256 checksum
How to use checksums
fd703a2092d336b0b3e731836448b61892cf9c3ac14822661406ea7bd20a9e8e
BLAKE2b-256 checksum
How to use checksums
51cc9e12f98e67dba51d4db50371c57c722e4df852e3aa6fc8f96b6bbd67137b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.5

Release files / aiar_rag-0.2.5-py3-none-any.whl

Download URL aiar_rag-0.2.5-py3-none-any.whl
Size 710.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
acd8de6208a56263407a70908f89a8ff8a82a1651b011a6accd0273d235c26d2
BLAKE2b-256 checksum
How to use checksums
37aa9bcb353d29d88a14709fd2044ce0661c56b571fd7a66b65493cf6dba6214
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.5

Release history Release notifications | RSS feed

This release

0.2.5 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page