Skip to main content

arara-rag

Portuguese-first retrieval that runs entirely on CPU. Chunking, dense and lexical retrieval, rank fusion and reranking — numpy only. No PyTorch, no ONNX Runtime, no FAISS. The whole install is 200 MB.

pip install arara-rag
from arara_rag import Arara

arara = Arara(path="./indice")            # out-of-core and persistent
arara.add_documents(
    {"lei_1234": open("lei.txt").read()},
    metadata={"ano": 2024, "tipo": "lei", "uf": "BR"},
)
hits = arara.search("qual a alíquota?", top_k=5, where={"ano": {"$gte": 2020}})
print(hits[0].doc_id, arara.resolve(hits[0]))     # exact source span

What's in it

Stage Component Size
Chunking tinyzchunk — tokenizer-free, distilled from an LLM teacher 2.1 MB
Dense static-nomic-384-pten-v2 — Model2Vec static embeddings 62 MB
Lexical BM25 over a numpy inverted index —
Rerank CXM25 — PT-BR lexical scoring bundled
Fusion Reciprocal Rank Fusion —

An index can be out-of-core: vectors live in a memory-mapped file and documents, metadata and offsets in SQLite, so the index is about 1 KB per document and cold pages can be evicted by the OS instead of being pinned on the heap. Same API either way — Arara() keeps everything in memory.

Speed and memory

One CPU core, no GPU. Measured end to end with python -m bench.profile.

documents index build query p50 query p95 index size peak RSS to serve
1,000 3.9 s 0.44 ms 0.45 ms 1.8 MB 472 MB
10,000 8.3 s 0.82 ms 3.3 ms 18 MB 479 MB
50,000 29 s 6.3 ms 9.6 ms 91 MB 493 MB
  • ~1,600–1,900 documents/second to chunk, embed, tokenise and index — chunking and embedding are per-document, so they run across processes. A single process manages ~280/second, and the out-of-core build costs the same as the in-memory one to within 10%.
  • ~1.8 KB per document of index: 1.5 KB of vectors plus BM25 postings.
  • Query latency scales with corpus size because both retrievers score the whole corpus per query — that is what makes the ranking exact rather than approximate.

Memory to serve does not scale with the corpus

Fifty times the documents costs 21 MB more, not fifty times more: a 50,000-document index serves in 493 MB, a 1,000-document one in 472. Everything corpus-shaped is memory-mapped, read a block at a time, and dropped again with madvise(MADV_DONTNEED) as soon as the block has been scored: the dense vectors, the BM25 postings, and the vocabulary (a sorted term blob read by binary search, because a Python dict of terms would cost ~140 bytes each).

What is left is a fixed floor of about 445 MB, measured by loading the stack one piece at a time:

MB
Python + this package + numpy 35
quantized embedding table (safetensors) 52
XLM-R tokenizer tables (276,214 tokens) 316

The tokenizer is the whole story, and it is not something the corpus can change. On top of that floor, max_ram_mb sets a hard ceiling on the whole process: the scan block is re-derived from the live footprint before every query, so the process stops short of the budget rather than growing into it.

arara = Arara(path="./indice", max_ram_mb=640)   # never exceeds 640 MB RSS

Squeeze it below roughly 500 MB and the scan has to fall back to its minimum block, which costs query speed but not correctness — the ranking is identical at every setting.

Speed and memory

Retrieval quality

MTEB-BR, the Brazilian Portuguese benchmark with a public leaderboard. nDCG@10, fixed-window chunking. Metrics are computed by bench/metrics.py, which bench/validate_metrics.py checks against pytrec_eval to 0.0e+00, and against scikit-learn on binary relevance.

Both dense backends are shown, because the choice matters more than any parameter in the stack. static is the default; nanoE5 is opt-in (see below).

Task docs rel./query static dense nanoE5 dense static hybrid nanoE5 hybrid lexical
BRTaxQAR (capped) 478 2.92 0.2934 0.3423 0.3486 0.4180 0.4051
FaQuADIR 244 1.0 0.7139 0.8314 0.8304 0.8906 0.8961
FaqBacenRetrieval 1,673 1.0 0.3744 0.5858 0.4526 0.5659 0.4881
JurisTCU 16,045 15.0 0.3887 0.4906 0.4890 0.5685 0.5378
Quati 50,000 38.66 0.3268 not run* 0.4046 not run* 0.4067

* Quati's nanoE5 index did not finish in the time available — it is feasible (~15 minutes of encoding), but this host was carrying load from other tenants and three attempts stalled past 45 minutes. Everything else in this table was run twice with identical results.

The last column is BM25 alone. Note what changes with a real encoder: hybrid beats lexical on three of the four rows where both were measured, and ties it on FaQuADIR. With the static model lexical won every row, so the argument for hybrid retrieval was not visible in these numbers until nanoE5 was added.

Dense and hybrid per backend against BM25

Read rel./query before comparing across rows. nDCG@10 measures very different things on these tasks, and that is a property of the benchmarks rather than of the retrieval:

  • FaQuADIR and FaqBacen have exactly one relevant document per query. There nDCG@10 is a transform of the rank of that one document, so 0.90 means it is usually first and 1.0 is the ceiling.
  • Quati has 38.66 relevant documents per query. With ten slots, a perfect score requires all ten to be relevant, so 0.41 means roughly four of the top ten are relevant — much closer to precision at 10 than the FaQuADIR number is.
  • JurisTCU (15.0) and BRTaxQAR (2.92) sit in between.

A retriever scoring 0.33 on Quati is therefore not "worse" than one scoring 0.71 on FaQuADIR — the numbers are not comparable across rows, only within one.

Choosing a dense encoder

Two backends ship. The default is a Model2Vec lookup table; nanoE5.c is an opt-in alternative — a 4-bit multilingual-e5-small in C, with no PyTorch, no ONNX and no BLAS. It is a real transformer forward pass, so it is slower, and it is a separate dependency:

pip install "arara-rag[nanoe5]"
Arara(dense_backend="nanoe5")                              # or "static", the default
Arara(dense_backend="nanoe5", dense_variant="original")    # English-first build

It is meaningfully better: dense retrieval improves by 0.05–0.21 nDCG@10 — on FaqBacen that is a larger jump than the whole distance from BM25 to the leaderboard median. The comparison table above is the full picture.

The cost is indexing speed, and it tracks chunk length because nanoE5 windows anything past 512 tokens. On a fixed corpus of 300 passages of ~1.1 kB:

static nanoE5
index build 2.5 s (121 docs/s) 14.4 s (21 docs/s)
query p50 0.5 ms 13 ms

Longer chunks cost more than proportionally — a corpus of 2.5 kB chunks is far slower again — while queries are one short sequence either way. nanoE5's query latency also has a longer tail than the static model's, so prefer static where p99 matters more than recall.

CXM25 reranking on top of the hybrid adds +0.018 to +0.077 nDCG@10 across these tasks for 1–6 ms per query (FaQuADIR: 0.8304 → 0.9078, BRTaxQAR full documents: 0.4801 → 0.5091).

Why chunking matters most

Legal documents in BR-TaxQA-R average 32,000 characters and reach 1.17M. MTEB-BR truncates them at 32k because transformer encoders cannot fit more. arara chunks, so it indexes the whole statute.

Configuration chunks nDCG@10 R@100
capped at 32k, one vector per doc (the leaderboard's setting) 478 0.1496 0.4351
capped at 32k, fixed windows 2,552 0.2934 0.6300
capped at 32k, tinyzchunk 23,319 0.3088 0.6225
full documents, fixed windows 6,439 0.4041 0.7497
full documents, paragraph splits 6,439 0.4041 0.7497
full documents, tinyzchunk 60,927 0.4287 0.7486
full documents, tinyzchunk + BM25 60,927 0.4801 0.8496
full documents, + CXM25 rerank 60,927 0.5091 0.8102

Chunking a legal corpus beats truncating it by 3.4×

This ablation is measured with the static backend. nanoE5 is not run here, and the reason is a property of that model rather than a gap in the table: it windows anything past 512 tokens, so BR-TaxQA's 2.5 kB chunks cost it about 0.9 chunks/second. The capped corpus alone would take ~17 minutes and the full-document, tinyzchunk configuration — 60,927 such chunks — would take most of a day. The comparison that matters for nanoE5 is in Retrieval quality, where a real encoder changes the conclusion about hybrid retrieval.

Against the leaderboard

Best configuration per backend, and how much of the field each beats:

Task static best nanoE5 best best on the leaderboard leaderboard median
FaQuADIR 0.9078 (hybrid+CXM25) 0.8906 (hybrid) voyage-context-4 (0.8738) 0.7689
BRTaxQAR 0.4051 (lexical) 0.4180 (hybrid) voyage-finance-2 (0.4499) 0.2723
JurisTCU 0.5378 (lexical) 0.5685 (hybrid) llama-embed-nemotron-8b (0.6805) 0.5436
FaqBacenRetrieval 0.4881 (lexical) 0.5858 (dense) codestral-embed (0.8262) 0.6516
Quati 0.4211 (hybrid+CXM25) not run voyage-context-4 (0.6901) 0.5569
Task static dense static hybrid static lexical nanoE5 dense nanoE5 hybrid
FaQuADIR 31% 80% 100% 81% 100%
BRTaxQAR 56% 76% 95% 71% 97%
JurisTCU 25% 34% 47% 35% 62%
FaqBacenRetrieval 18% 27% 30% 33% 30%
Quati 22% 29% 29% - -

Percentages are the share of the 95-96 published models each beats. On FaQuADIR arara outranks every model on the board, and on BRTaxQAR the nanoE5 hybrid beats 92 of 95. The static model's dense retrieval is the weak cell in every row — 18% on FaqBacen, 25% on JurisTCU — and nanoE5 lifts both.

The honest caveat: the leaderboard evaluates embedding models, and there is no BM25 entry on it. arara's strongest static modes are lexical, and lexical retrieval is simply very good on short, high-overlap PT-BR documents — part of that gap is a missing baseline on their side.

arara against the MTEB-BR leaderboard

Reranking

MTEB-BR reranking hands you a fixed candidate list and scores only the order (MAP@1000), so identity is the baseline the benchmark ships with.

Task identity static dense nanoE5 dense static hybrid nanoE5 hybrid CXM25
QuatiReranking 0.2839 0.2798 0.5003 0.3066 0.4300 0.3100
JurisTCUReranking 0.4150 0.3609 0.4616 0.4129 0.4698 0.4845

This is the sharpest version of the same story. On QuatiReranking the static dense model regresses on the candidate order it was given (0.2839 → 0.2798), while nanoE5 dense improves it by +0.22 MAP@1000 and beats every static mode, CXM25 included. Lexical reranking is unaffected by the backend, as it must be.

Out-of-core, metadata, CRUD

An index is read far more than it is written, so deletes are tombstones and freed slots are recycled on the next write.

arara = Arara(path="./indice", max_chunk_chars=2000, max_ram_mb=512)

arara.add_documents(docs, metadata={"ano": 2024})        # insert / replace
arara.update_metadata("lei_1234", {"revisado": True})    # no re-embedding
arara.delete_document("lei_1234")                        # tombstone + slot reuse
arara.get_document("lei_1234")                           # (text, metadata)
arara.compact()                                          # reclaim file space

arara.search(q, where={"tipo": {"$in": ["lei", "decreto"]}, "ano": {"$gte": 2020}})
arara.search(q, where={"$or": [{"uf": "SP"}, {"uf": "RJ"}]})

max_ram_mb is optional and only meaningful with path=; it bounds the index data a query may hold resident, not the ~445 MB interpreter and model floor (see above). Leave it unset to use the default 24 MB scan block.

Supported per field: $eq (bare value), $ne, $gt, $gte, $lt, $lte, $in, $nin, $exists, $contains, $startswith, $endswith; plus top-level $and / $or. Field names are validated and values are bound as SQL parameters, so a filter cannot inject SQL.

Guarantees

Enforced by 120 tests, not asserted in prose:

  • every chunk is an exact substring of the canonical document, ordered and non-overlapping, with only whitespace between chunks — nothing is dropped;
  • no chunk exceeds max_chunk_chars, including on a 24,000-character line;
  • CRLF and LF inputs chunk identically and offsets still resolve;
  • the in-memory and on-disk paths return identical rankings;
  • capping the scan budget with max_ram_mb never changes the ranking, only how much of the index is resident at once;
  • importing the package never imports torch or onnxruntime.

Reproduce

python -m venv .venv && .venv/bin/pip install -e ".[bench,validate,dev]"
python -m pytest tests/                 # 120 tests
python bench/validate_metrics.py        # metrics vs pytrec_eval
./bench/run_all.sh                      # every suite -> bench/results/
python -m bench.profile                 # speed and memory -> docs/scaling.png
python -m bench.ramcheck run 10000 40000 100000   # serving RSS vs corpus size
python -m bench.charts                  # regenerate the figures
python -m bench.leaderboard             # compare against MTEB-BR

Layout

arara_rag/
  chunk.py      chunking and the losslessness contract
  dense.py      static encoder
  lexical.py    BM25 inverted index + CXM25 reranker
  store.py      memory-mapped vectors, SQLite catalog, filters
  pipeline.py   Arara: add / search / rerank / CRUD
bench/          task loaders, metrics, suites, profiling, charts
tests/          120 contract and correctness tests
space/          Gradio demo

Limitations

  • PT-BR and English. The tokenizer, stemmer and stopwords are Portuguese.
  • The dense model is small and static; lexical retrieval carries the stack on short, high-overlap documents.
  • Per-chunk bookkeeping stays resident (16 bytes/chunk); vectors, text and metadata do not.
  • CXM25 reranking is ~71 µs/document, so it runs over a candidate set.
  • Index build forks worker processes to parallelise chunking and embedding. Set workers=1 where forking is unsafe or unavailable; the result is byte-identical, only slower.

License

Apache-2.0.

Metadata

Release files for arara-rag 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for arara-rag 0.6.0
File Size Uploaded
arara_rag-0.6.0.tar.gz 57.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for arara-rag 0.6.0
File Interpreter ABI Platform
arara_rag-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 100.0 kB

Release files / arara_rag-0.6.0.tar.gz

Download URL arara_rag-0.6.0.tar.gz
Size 57.9 kB
Tags Source
SHA-256 checksum
How to use checksums
a92bab13dab19d96799144a374cfa2aff5531262f2b5b3806130a1f13c72252f
BLAKE2b-256 checksum
How to use checksums
0f45f0f4d57ae930113a107cfbffe2baad2e033ee44c5843a118923370d75814
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.7

Release files / arara_rag-0.6.0-py3-none-any.whl

Download URL arara_rag-0.6.0-py3-none-any.whl
Size 42.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
aae23b69adf17a477ff28a8b86c0b38137c9db82c7c351727ab06bd5d6f98ce2
BLAKE2b-256 checksum
How to use checksums
2d611dcd59b6e40e34b25af33c8625267fb0b6457f46b1d2d68646d04e29d6e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.7

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page