whoosh-compat
A standalone, typed, Python 3.11+ library that parses the
Whoosh query language into a
backend-neutral AST and emits programmatically constructed
tantivy.Query objects: never
via tantivy's own string query parser.
It exists so that applications which used to build Whoosh query objects
directly (or hand-translate Whoosh-style query strings into another engine's
query language) can keep their existing, lenient, natural-language search
syntax while running on tantivy.
The motivating case is paperless-ngx,
which migrated its search backend from Whoosh to tantivy but kept Whoosh's
query syntax as its user-facing search language. See
paperless-ngx#13568
for a concrete example of a query (title:202[0-3]*, a bracket-class
wildcard) that a naive string-translation layer gets wrong.
Installation
pip install whoosh-compat[tantivy]
The core package (pip install whoosh-compat) depends only on
python-dateutil. The tantivy extra pulls in tantivy-py, which is
required to use whoosh_compat.emitters.tantivy_. The AST itself
(whoosh_compat.ast) and the parser (whoosh_compat.parse) have no tantivy
dependency, so a future backend (e.g. Meilisearch) can reuse them.
Usage
import tantivy
import whoosh_compat as wc
from whoosh_compat.emitters.tantivy_ import emit
# 1. Describe the fields the parser and emitter need to know about.
registry = wc.FieldRegistry(
[
wc.FieldSpec("content", wc.FieldKind.TEXT, analyzer=str.split),
wc.FieldSpec("tag", wc.FieldKind.KEYWORD, comma_values=True, analyzer=str.split),
wc.FieldSpec("created", wc.FieldKind.DATE, date_only=True, fast=True),
]
)
# 2. Parse a whoosh-style query string into a backend-neutral AST.
result = wc.parse(
"tag:steuer AND created:[2020 TO 2020]",
registry=registry,
default_fields=["content"],
)
result.ast # normalized AST root
result.diagnostics # tuple[Diagnostic, ...], e.g. invalid dates
# 3. Emit a tantivy.Query against a real index and search it.
query = emit(result.ast, index=index, registry=registry)
searcher.search(query, limit=10)
FieldSpec.analyzer is how index-time tokenization parity is achieved: it's
a plain Callable[[str], list[str]] the host supplies (e.g.
tantivy.TextAnalyzer.analyze), called at emit time on Term/Phrase
text. tantivy's term_query does not tokenize on its own, so skipping this
step means a query for "Invoices" silently fails to match an index that
stored the lowercased token "invoices".
See ARCHITECTURE.md for the full FieldSpec/
FieldRegistry shape and how a query string gets from string to
tantivy.Query.
API Stability
The public API boundary, as of 0.1.0:
Stable API (guaranteed across minor and patch releases):
whoosh_compat.parse()andwhoosh_compat.ParseResultwhoosh_compat.astmodule (the backend-neutral query tree), includingwhoosh_compat.free_text_tokens()(also exported at top level)whoosh_compat.fieldsmodule (FieldSpec,FieldRegistry,FieldKind, etc.)whoosh_compat.errorsmodule (exception types)whoosh_compat.emitters.tantivy_.emit()function and theEmitterprotocol inwhoosh_compat.emitters.base
Internal / not stable (usable, but subject to change without notice between versions):
whoosh_compat.parser.*: the forked whoosh tagger and filter pipeline. The parser is a fork of whoosh's own query parser, kept close to upstream so it stays diffable and easy to maintain. Because it tracks a third-party codebase, its internals and behavior may change between whoosh-compat releases, even minor ones.
The parser.* exemption forecloses nothing: it can always be promoted to stable later. In the meantime, if you import directly from whoosh_compat.parser, your code may need updates on whoosh-compat releases.
Module naming: why emitters.tantivy_?
The emitters.tantivy_ module uses a trailing underscore to signal that this is a backend-specific module. This naming is deliberate and permanent at the 0.1.0 release. While a future version could add a lazy re-export (e.g., emitters.tantivy without the underscore) to provide an alternative import path, the canonical import is and will remain from whoosh_compat.emitters.tantivy_ import emit. The underscore also reinforces that tantivy is an optional dependency: the emitter won't be imported unless you ask for it.
The host contract
A host embedding this library (mapping failures to an HTTP 400, for example) needs to check for two independent failure modes, not one:
ParseResult.diagnosticsis non-empty.parse()never raises for bad query input; a malformed date, an out-of-domain number, or other field-kind-specific problem becomes aDiagnosticplus anErrorLeafin the tree instead. Callingemit()on a tree containing anErrorLeafraisesQueryError.emit()raisesQueryError. This can happen even whendiagnosticsis empty: some query shapes parse cleanly but have no way to execute against tantivy today. The canonical example is a text-field range,title:[a TO b]: whoosh supported this, but tantivy-py has no programmatic text-range API (DIVERGENCES.mdentry 5), so the query parses withdiagnostics == ()and only fails onceemit()is called. There is a single exception type now:QueryErroralways carries aDiagnostic(err.diagnostic) describing why.
An empty diagnostics tuple does not, by itself, mean emitting is safe.
Both checks matter; a host that only looks at diagnostics will still see
an uncaught QueryError bubble up for shapes like the one above.
The one exception parse() itself can raise is QueryParserError, and it
never means the query was bad. It means a defect in this library: the parse
pipeline is wrapped in a backstop that converts any unexpected exception into
that type, chaining the original as __cause__, so a host routes it to a
monitorable 500 instead of seeing (for example) a bare RecursionError from a
pathologically nested query. It is deliberately not a Diagnostic: reporting
an unknown internal failure as a 400 would blame the user for a bug on this
side and hide it from monitoring. (Misconfiguration passed to parse() --
an empty or unknown default_fields, a naive basedate -- still raises
ValueError eagerly, as documented on the function.)
Branch on Diagnostic.kind and Diagnostic.cause; treat message as
log output, never parse it. Both ParseResult.diagnostics entries and a
caught QueryError's .diagnostic are structured records: kind is a
stable DiagnosticKind member a host can switch on (for example, mapping
BAD_DATE to a typed InvalidDateQuery), and cause is a coarser Cause
a host can use for routing without knowing every DiagnosticKind:
Cause |
Meaning | Typical host response |
|---|---|---|
INVALID_INPUT |
The query text itself is malformed | HTTP 400 |
UNSUPPORTED |
The query is well-formed but this backend can't run it | HTTP 400 |
MISCONFIGURED |
The registry/schema setup is wrong | Operator alert and HTTP 400 |
INTERNAL |
The AST violates an invariant parse() would never produce |
HTTP 500 |
MISCONFIGURED is the one cause that is two responses rather than one. It
means the registry and the index schema disagree, which only an operator can
fix, so it must raise an alert. But every MISCONFIGURED kind is reachable
from ordinary query text (notes.user:* for EXISTS_REQUIRES_FAST, and any
query naming a field the registry knows and the index schema does not for
SCHEMA_FIELD_MISSING), so a request is waiting on an answer that the alert
does not provide. The query cannot run whether or not anyone reads the
alert, and reporting it as a 500 would claim a defect in this library that
isn't there, so the request gets a 400.
SCHEMA_FIELD_MISSING is reported uniformly by every leaf that queries a
resolved field (term, phrase, prefix, wildcard, numeric and date range,
bare-* existence, and JSON subpaths): the drift is a property of the
field, not of the spelling that reaches it, so content:x and content:x*
never land on opposite sides of the 400/500 line for the same broken
deployment. Only the confirmed missing-field condition is reclassified; any
other refusal from tantivy-py remains BACKEND_REJECTED/INTERNAL, so a
genuine defect in this library is never hidden behind a 400.
Diagnostic.message (and the QueryError exception message, which is the
same string) carries no stability guarantee and may reword without notice.
Everything a host needs to act on is on the record's own fields instead: a
DIVERGENCES.md entry number on divergence, the field and its kind on
field/field_kind, the offending literal on raw_value, the span in
the query string on startchar/endchar, and, where a single concrete
rewrite of the query text would work, that rewrite on suggestion (see
"Adopting the library" below for how to apply it, and why it is a separate
field rather than a new DiagnosticKind).
Treat the message as developer/log output: a host showing errors to end
users should build its own copy from kind/cause, not display or parse
message; the paperless-ngx integration does exactly this.
A Diagnostic's severity is fatal-only, and always will be: there is no
severity field, and none is planned. Any diagnostics present means the
query cannot be emitted, full stop; there is no "warning" tier to weigh
differently. A future informational-only signal (for example, reporting
that a zero-token term was silently dropped during analysis) would use a
separate channel, never ParseResult.diagnostics.
Free-text tokens for secondary clauses
whoosh_compat.free_text_tokens(node, registry=..., fields=...) answers
"which plain words does this query search for?" for hosts that blend a
secondary text clause alongside the emitted query: the motivating case is
paperless-ngx's fuzzy-matching blend, which re-parses a word string
through tantivy's own query parser and must never receive whoosh grammar.
It returns the analyzed Term/Phrase tokens on the requested
TEXT/KEYWORD fields, deduplicated in first-appearance order, with the
subtle rules handled here rather than in each host: negated subtrees
(NOT x, ANDNOT's negative side) contribute nothing, patterns and
ranges contribute nothing, and a word the multifield expansion copied
onto several default fields counts once. Tokens are the field analyzer's
output verbatim, never re-split; see the function's docstring for the
full contract.
Pass the tree as parsed (ParseResult.ast). The negation rule is why:
analyze() deliberately collapses an AndNot whose positive side
analyzed to nothing, leaving the negative side standing alone as an
ordinary positive node (DIVERGENCES.md entry 23), so a tree you analyzed
yourself before calling no longer records what the user excluded, and no
walk can recover it.
Pass analyzed=False to get the raw text each contributing leaf was
parsed from instead of the analyzer's output. Do that whenever the tokens
are going back into a parser that will analyze them again: analysis is
not generally idempotent (a stemmer maps universities to univers and
then univers to univ), so re-analyzing analyzed output searches for
something the index does not contain. In that mode the analyzer is never
consulted. Which nodes contribute is structural (negation, patterns,
kinds and dedupe never vary by mode) with one exception, the zero-token
leaf in the first row below; the other two rows differ in text only:
| query | analyzed=True |
analyzed=False |
|---|---|---|
the (a stopword) |
() |
('the',) |
"tax reports" |
('tax', 'report') |
('tax reports',) |
alpha-beta (analyzer splits it) |
('alpha', 'beta') |
('alpha-beta',) |
An all-stopword leaf analyzes to nothing, so it contributes no entry at all in analyzed mode, while unanalyzed mode reports its raw text: that is a node present in one mode and absent from the other, not two spellings of one node. Deciding membership by the analyzer while refusing its output would be a half-analysis this mode's contract denies; the re-parse downstream applies its own stopword list, in its own index's terms.
So an entry in this mode can contain whitespace and punctuation, and the
number of entries is the number of contributing leaves rather than of
words. It can also contain characters a re-parse would read as grammar
(a colon, a bracket) even though the query grammar around them is gone:
quote or escape before re-parsing, and note that a whole-token filter
(\w+) over these entries drops every hyphenated, dotted or
quoted-phrase term outright, because the text is untokenized.
Cap query length at the host boundary. Parse time is linear in query
length on every shape measured, the adversarial ones included: a long run
of word characters with no : (CJK text, a long token, a dotted string),
unclosed range brackets, and unmatched single quotes. Each used to be
super-linear, because a tagger's regex rescanned the rest of the input
from every position it was tried at, and the worst, an unclosed [ with a
to after it, grew with the cube of the length (a 4KB query cost about 30
seconds). Those taggers now work out where their expression cannot match
and skip the scan, giving exactly the same parse: measured on one
developer machine, 16KB of any of those shapes parses in about a second
and 64KB in a few seconds. The parser's own nesting-depth cap bounds
recursion, not CPU time, and linear is not free, so a host accepting
untrusted query strings should still enforce its own length limit (a few
KB comfortably covers any human-written query) before calling parse().
Hand-building a Fuzzy node for a caller-side companion clause
whoosh_compat.ast.Fuzzy is a leaf node this library never produces
itself: there is no ~ grammar registered in this library's parser
plugin set (real Whoosh has one; it is deliberately not carried over, see
"Not carried over from Whoosh" below), so a fuzzy (edit-distance) match is
only ever hand-built, the same way a caller may already hand-build a tree
containing ast.Nothing()/ast.Every(). A tree containing one can go
straight to emit(). For a companion clause around each word of a parsed
query, place it with the rewrite_leaf hook (passed to emit())
instead, returning Or(leaf, ast.Fuzzy(...)), which keeps the word's own
analysis exactly
as it was (see "Rewriting leaves before emit" below). The hook sees
leaf.text as the raw, unanalyzed query text, though: Fuzzy(text=str(leaf.text))
on alpha-beta is one fragment through pattern_normalizer, not two
tokens, and will not match a tokenized index. Split the text the way the
field's analyzer would first, and build one Fuzzy per token.
from whoosh_compat import ast
from whoosh_compat.fields import FieldRef
fuzzy_leaf = ast.Fuzzy(field=FieldRef("content"), text="tokyo", distance=1)
Unlike Term/Phrase/Prefix/Wildcard, field is required, not
optional: there is no "expand across default search fields" behavior for
a fuzzy leaf. A caller wanting fuzzy matching on several fields builds an
Or of several explicitly fielded Fuzzy nodes. Give FieldRef the
field's canonical name, not an alias: FieldRegistry.resolve() does
accept an alias here (it currently has no way to tell a hand-built
FieldRef from a parser-produced one), but the AST's own invariant is
that a FieldRef already carries the canonical name, and
FieldRegistry.make_ref is where a parser-typed alias is meant to be
canonicalized to it.
text is matched via FieldSpec.pattern_normalizer (the same
fragment-level, non-tokenizing normalization seam Wildcard/Prefix
already use, see "The analyzer / pattern_normalizer seam" below), never
the full analyzer: analysis never rewrites a Fuzzy leaf, even when one
passes through ast.analyze() inside a rewrite_leaf replacement, so if
a field's normalizer offers several candidate forms, every form is tried
and a matching document is scored by its best-matching form, not once per
matching form. A field with no pattern_normalizer configured
uses text exactly as given. A host whose field lowercases at index time
should configure a pattern_normalizer, or a case difference alone
consumes the edit-distance budget with nothing left for the typo it was
meant to tolerate: at distance=1, Fuzzy(text="Tokyo", distance=1)
against an indexed tokyo spends its whole budget matching the case
difference and has none left for an actual misspelling, and at
distance=0 it matches nothing at all.
A blank (empty or whitespace-only) text matches nothing, and so does a
text the pattern_normalizer reduces to blank (say, punctuation
only): blank forms are dropped from the alternatives, and a node left
with none matches nothing. It is not an error, since parse() itself
produces empty terms from ordinary input like '', and a host mirroring
a parsed Term into a Fuzzy should not have to filter them out.
Passed through to tantivy, a blank term would match every one-character
term at distance=1 and every term in the field with prefix=True.
Short words need the host's own care with prefix=True, though: any
text no longer than distance matches every term in the field, since
the empty start of every term is within distance edits of it.
Fuzzy(text="x", distance=1, prefix=True) matches everything. A host
building a companion clause from the user's words should skip words that
short, rather than rely on this library to reject them, since they are
legitimate input.
Only TEXT and KEYWORD fields are supported. Every other FieldKind,
including a JSON field, subpath or bare, fails at emit() time with
AST_KIND_NOT_IMPLEMENTED: tantivy-py's fuzzy query API (as of 0.26.0)
has no way to scope a match to one JSON subpath, and a bare JSON field
wants a JSON value argument, not a term string, so neither shape is
supported. The limit is in tantivy-py's binding; tantivy's own fuzzy
query can match within a JSON path.
distance (default 1) and prefix (default False) map directly onto
tantivy.Query.fuzzy_term_query's own parameters. distance must be
the integer 0, 1 or 2, the only distances tantivy can run a fuzzy
search with; any other value, including a bool, float or str,
fails at emit() time with AST_BAD_NUMBER. Without that check,
tantivy-py would accept any distance up to 255 when building the query
and only reject it once searcher.search() runs, as a bare ValueError
the host would have to catch itself.
A text that is not a str, or a prefix that is not a bool, fails
with AST_INVALID_SHAPE.
prefix=True matches every indexed term that starts with something
within distance edits of text: Fuzzy(text="tok", distance=0, prefix=True) matches tok, tokyo and tokio. It widens the match.
Whoosh's FuzzyTerm has a similarly named prefixlength that does the
opposite, narrowing the match by requiring the first N characters to
match exactly; Fuzzy has no equivalent of it.
transposition_cost_one is not exposed as a field on Fuzzy; it is
always True (tantivy's own default).
Rewriting leaves before emit
A host that widens a parsed query, for example searching an internal
companion field alongside each Term or Phrase word in it, does it
through the rewrite_leaf hook instead of walking the tree itself. Pass
the hook to emit(), which runs it inside its own analysis pass. Pattern
leaves are not passed to the hook, so a wildcard or prefix word such as
invoi* gets no companion:
import whoosh_compat as wc
from whoosh_compat import ast
from whoosh_compat.emitters.tantivy_ import emit
# emit_registry is registry plus the companion field, declared with
# multitoken=wc.Multitoken.AND so all of its tokens must match:
# wc.FieldSpec("content_grams", wc.FieldKind.TEXT, analyzer=...,
# multitoken=wc.Multitoken.AND)
def widen(leaf: ast.Term | ast.Phrase) -> ast.Node:
if leaf.field is None or leaf.field.name != "content":
return leaf
companion = ast.Term(field=wc.FieldRef("content_grams"), text=str(leaf.text))
return ast.Or(children=(leaf, companion))
result = wc.parse(q, registry=registry, default_fields=["content"])
query = emit(result.ast, index=index, registry=emit_registry, rewrite_leaf=widen)
ast.analyze() takes the same keyword, and
emit(ast.analyze(result.ast, emit_registry, rewrite_leaf=widen), ...)
builds the same query, for a host that wants the analyzed tree itself. It
costs a second analysis pass inside emit(), so pass the hook to emit()
when the query is all you need.
A walk of your own gets two things wrong that the hook gets right, because the hook runs inside analysis's own pass:
- Multi-token context. Wrapping a leaf in a new
Orchanges its enclosing group, and aMultitoken.DEFAULTterm that the analyzer splits into several tokens is combined according to that group (DIVERGENCES.md entry 15). Socontent:alpha-beta AND reportwould loosen from requiring both tokens to accepting either. - Zero-token drops. Analyzing the leaf yourself and putting the result
back does not fix that: a leaf that analyzes to zero tokens then becomes
a pre-existing empty operand, which empties an
AND(entry 27), instead of dropping out of it (entry 23).
What the hook sees and what its answer means:
- It is called with every
TermandPhrasein the normalized tree, of any field kind, never with any other node type, and never with a node inside a replacement it returned. It is called once per leaf object and AND/OR context: one leaf object placed at two positions under the same context gets one call and one result. Call order is unspecified. - It is called for a leaf under a negation too:
NOT x, or the negative side of anAndNot. A companion placed there widens what the negation excludes rather than what it matches, since negatingOr(leaf, companion)excludes anything either side matches. WithOr(leaf, companion, fuzzy),NOT tokyowould also exclude every document with a fuzzy match such astokio. For already-normalized input (everyparse()result is), the leaf each call receives is the input tree's own object, so a host that wants to skip a negated leaf can pre-scan the tree it passes in for leaves under a negation and compare by identity (is). - Returning the leaf itself changes nothing. Return
ast.Nothing()to remove the leaf: it drops out of its group exactly as a stopword would. - Anything else replaces the leaf and is analyzed in the leaf's position.
Put the leaf object itself into the replacement, not a copy. Wherever it
appears there, it stands for its own analysis in its original position,
so
Or(leaf, companion)keepsalpha-betarequiring both tokens. A copy is analyzed as a new leaf inside yourOrand loosens. - A replacement that ends up empty, for example
Or(leaf, companion)where both sides analyze to nothing, drops out like a stopword. - Other leaves in the replacement, such as the companion, are analyzed as
new leaves. A companion on a
Multitoken.DEFAULTfield that the analyzer splits into several tokens takes its combination from its own group in the replacement, usually yourOr, so any one of its tokens matches. Declare the companion field with an explicitmultitoken(for exampleMultitoken.AND) when all of them must match. - Companion leaves are analyzed with the registry you pass alongside the
hook. Parse with the registry users may address, and emit (or analyze)
with one that also carries the companion fields.
parse()reads an unknown field prefix as literal text, so an internal field stays unreachable from query text. - Combine several rewrites (say a companion field and a fuzzy clause) into
one hook that returns
Or(leaf, companion, fuzzy). Don't apply a hook twice, for example toanalyze()and then again toemit(): the second pass would see the split tokens and companions, not the words the user typed. A treeanalyze()already returned is fully analyzed, andemit()without a hook changes nothing about it.
Errors through emit(): the hook runs inside emit()'s input stage, so
its errors are handled the way a field analyzer's are. A ValueError,
TypeError, AttributeError, NotImplementedError or RecursionError
(or a subclass) from the hook or an analyzer, and a hook return value that
is not an ast.Node, become a QueryError with AST_INVALID_SHAPE, whose
cause is INTERNAL (a 500, not the user's fault), with the original
exception chained as its context. Any other exception reaches you
unchanged. A replacement is checked like any hand-built tree, so a
malformed one fails the way the same shape passed to emit() directly
would. Errors through analyze(): an exception raised by the hook or by a
field analyzer reaches you unchanged, and a hook return value that is not
an ast.Node raises TypeError. Treat anything other than QueryError
from either call as an internal error, the way emit()'s own
AST_INVALID_SHAPE is one.
Adopting the library: sweep stored queries first
A host switching to this library from real Whoosh usually carries a body of
stored queries written against the old engine: saved views, bookmarks,
scheduled searches. Some of those queries never worked the way their author
intended, and real Whoosh gave no sign of it. Parsed by the pinned oracle at
a base date in 2026, created:december 2019 resolves to a window over
December 2026 and searches for 2019 as a free-text word. Nobody who
saved that query was told anything was wrong.
Where this library rejects such a value instead (DIVERGENCES.md entry 61),
the stored query stops returning wrong documents and starts returning a
diagnostic. That is the intended improvement, but it lands on users who were
not aware they had a broken query, so it is worth doing before cutting over
rather than discovering it in production.
Parse every stored query and look at what comes back. Three outcomes matter:
-
No diagnostics,
emit()succeeds. Nothing to do. -
A diagnostic carrying a
suggestion. The unquoted multi-word date family is the case that has one today, except where the run pairs a time of day with a whole period (outcome 3).suggestionis the replacement text for that diagnostic's ownstartchar/endcharspan, so the host splices rather than re-deriving the rule. Apply one query's diagnostics in descendingstartcharorder, so rewriting one value does not shift the spans before it:result = whoosh_compat.parse(q, registry=..., default_fields=..., tz=..., basedate=...) out = q for d in sorted(result.diagnostics, key=lambda d: -(d.startchar or 0)): if d.suggestion is not None: out = out[: d.startchar] + d.suggestion + out[d.endchar :]
stored query rewritten created:december 2019created:"december 2019"created:2020 to 2021created:"2020 to 2021"created:previous month to nowcreated:"previous month to now"created:december 2019 OR added:2020 august 4created:"december 2019" OR added:"2020 august 4" -
A diagnostic with
suggestion is None. No single rewrite of the query text would work: a malformed date (created:20231340), a value the date grammar does not recognise at all (created:last week, where quoting does not help either), a time of day on a whole period (created:"this month 15:00", which needs a day named or the time dropped), a pattern on a numeric or BOOLEAN_EXISTS field (type_id:1*), or a shape with two possible fixes that mean different things (title:200[1-9]). These need a human, or a decision to drop the clause.
Branch on suggestion is not None, not on kind. The same kind carries a
suggestion for one query and not for another: BAD_DATE covers both
created:december 2019, which has a working quoted spelling, and
created:20231340, which has none. That is why the signal is a separate
field rather than a new DiagnosticKind, which would also have broken every
host already branching on BAD_DATE.
Re-parsing the rewritten query before storing it is still worth doing as a belt-and-braces check, but it is no longer what tells the two cases apart.
Supported query syntax
Parity target is Whoosh's intended grammar, not every Whoosh plugin.
See ARCHITECTURE.md for why the parser is a fork rather than a
reimplementation, and DIVERGENCES.md for every point where whoosh-compat's
behavior intentionally differs from real Whoosh.
| Syntax | Example | Notes |
|---|---|---|
| Bare terms, implicit AND | invoice total |
both terms required (Whoosh's own semantics) |
| Boolean operators (uppercase only) | a AND b, a OR b, NOT a, a ANDNOT b, a ANDMAYBE b, a REQUIRE b |
lowercase and/or/not are plain text, matching Whoosh's operator regexes; REQUIRE is infix like the others |
| Grouping | (a OR b) AND c |
|
| Fielded terms | title:invoice |
|
| Field aliases | type:invoice → document_type:invoice |
host-configured on FieldSpec.aliases |
| Phrases | "exact phrase", "a b"~2 |
slop follows Whoosh semantics: 1 = adjacent |
| Wildcards | inv*, inv?ice, 202[0-3]* |
glob syntax including bracket character classes |
| Prefix | inv* (no other wildcard chars) |
folds to a prefix query |
| Ranges | created:[2020 TO 2025], asn:[100 to] |
numeric and date; open-ended on either side |
| Boost | title:invoice^2.0 |
|
| Comma value lists | tag:foo,bar → tag:foo AND tag:bar |
per-field opt-in (FieldSpec.comma_values); tag:'foo,bar' quoting keeps it a single literal |
| Every / exists | *, *:*, title:* |
|
| Dates | created:2020, created:today, created:previous month, created:'previous month', created:now-7d |
full grammar: ISO/compact forms (a colon separates clock units only, see DIVERGENCES.md entry 63), natural-language keywords, relative offsets (now-7d, -1 week); the six multi-word keywords (previous week/month/quarter/year, this month/year) parse quoted or bare, an extension over Whoosh, where a value ends at the first space; a time of day written after one of them belongs to the date value too, and is rejected: a time of day on a whole month, year, week or quarter (created:"this month 3pm", created:"august 2026 15:00", created:'2026 23:59') names nothing, and is a BAD_DATE in every spelling, see DIVERGENCES.md entries 19 and 62. Every other unquoted multi-word date value is rejected with a BAD_DATE diagnostic naming the whole value rather than silently truncated to its first word: created:december 2019, created:2020 to 2021 and even created:previous month to now (a keyword phrase is joined first, then read as the start of a longer run) all error, and the quoted or bracketed spelling (created:"december 2019", created:[2020 TO 2021]) is the one that works, see entry 61 |
| RFC3339 datetimes | created:[2020-01-01T00:00:00Z TO 2020-06-01T00:00:00Z], created:'2020-01-01T00:00:00Z' |
T joins the separator class and a trailing Z is honored as the UTC designator (an absolute instant, not local time); an extension over Whoosh, which cannot parse these correctly (quoted T/Z values parse to nothing; range bounds collapse to their leading year), see DIVERGENCES.md entries 12 and 48-50. Quote it or bracket it: the bare unquoted spelling (created:2020-01-01T00:00:00Z) is split at its colons by the field-separator rule before any date parsing happens, and the half of it that survives is rejected as a bad date rather than searched as the shorter period it looks like, see entry 54 |
| JSON subpaths | notes.user:alice |
an extension with no equivalent in Whoosh itself (registered per-field via FieldSpec.subpaths) |
| Default JSON subpath | notes:alice → notes.user:alice |
per-field opt-in: the one subpath declared SubpathSpec(default=True) is what a bare mention of the field means, so a host never has to rewrite notes: in the raw query string. Without a default, a bare JSON field name stays unrecognized and demotes to text |
Not carried over from Whoosh (not currently implemented, kept cheap to add
via the forked plugin architecture): asn:>100 (GtLtPlugin), term~2
fuzzy matching (no parser syntax exists for it, but a caller can still get
fuzzy matching by hand-building an ast.Fuzzy node, either in a tree
passed to emit() or placed around a parsed word with the rewrite_leaf
hook, see "Hand-building a Fuzzy node for a caller-side
companion clause" above), r"regex" literal regex queries, SequencePlugin,
-foo/+foo as negation/requirement shorthand (in the whoosh grammar this
library targets, -foo was plain text whose analyzer typically dropped the
dash: NOT was the only negation operator), and free-date mode (implicit
date parsing in an unfielded-date context; the parser defaults to the
date-parsing plugin for fielded dates when the host calls parse() with a
date-aware registry instead).
Divergences from real Whoosh
AST-level divergence does not always mean result-level divergence.
whoosh-compat is tested at two separate layers. See
DIVERGENCES.md entry 16 for the full explanation, and
tests/emitter/test_acceptance_e2e.py's module docstring for worked
examples: a documented, real difference in the parsed AST between
whoosh-compat and real Whoosh (e.g. how a wildcard pattern's case-folding is
sequenced) can still produce the same final matched-document set, because
the divergence gets absorbed somewhere downstream (e.g. both sides' text
analyzers end up lowercasing the same way regardless). Read
DIVERGENCES.md for what's intentionally different and why; don't assume an
entry there implies a query result bug without checking whether it's one of
the entries called out as AST-only.
The analyzer / pattern_normalizer seam
Two separate callables on FieldSpec, deliberately not unified into one:
analyzer(Callable[[str], list[str]]): the full token-level chain (lowercase, ASCII-fold, stemming, stopword removal, whatever the host's index-time analyzer does) applied toTerm/Phrasequery text at emit time, so query tokens match what's actually in the index.pattern_normalizer(PatternNormalizer, i.e.Callable[[str], str | Sequence[str]]): a narrower, fragment-level transform applied to each literal segment of aWildcard/Prefixpattern, and, as a third consumer of the same seam, to aFuzzyleaf'stext. It never tokenizes and never drops a fragment; beyond that it is usually lowercase + ASCII-fold, and on a stemmed field it also offers the segment's stem. It may return one form of the segment (a barestr) or several alternatives (a sequence); the emitter matches a term satisfying any of them. Inside a bracket class it is additionally held to being character-level, because that is the one place a fragment is a single character and must stay one (see below).
These have to be different callables. The analyzer answers "what tokens does
this value become"; the pattern normalizer answers "what could this
fragment of a glob look like in the index", which is not the same question:
a fragment is usually not a word (inv*oices), it can never be split into
several tokens or dropped, and inside a bracket class it has to stay exactly
one character long.
The alternatives are what let a pattern reach a stemmed index without
giving up the spelling the user typed. English Snowball substitutes rather
than truncates, so neither form alone is enough: company stems to
compani (only the stem finds the indexed term) while copyright is its own
stem (only the typed run finds it). A host with a stemmed field returns both
forms and gets both documents; measured over a 4,977-word vocabulary, 3.5% of
words stem to something that is not a prefix of themselves, so no
"use the stem when ..." rule over a single string separates the two cases.
Each literal run alternates independently, so a many-run pattern costs the
sum of its alternatives, not the product.
Two properties of the seam survive that widening. Inside a bracket class the
normalizer is still applied one character at a time and only when it answers
with exactly one alternative exactly one character long (a class position
matches one character, and every offset in the glob translation is taken
against the body's length). And a normalizer is never asked to be
correct on a fragment that is not a word: stemming oices (from
inv*oices) is morphological nonsense, but as an added alternative it can
only widen the match, never move it, which is why alternatives replaced the
single-string form rather than joining it.
Both callables are checked against those return types at emit time. An
analyzer must return tokens of str (a list, or any iterable of them; a
bare str is rejected rather than split into characters), and a
pattern_normalizer must return a str or a sequence of str. Anything
else fails the query with AST_INVALID_SHAPE (cause INTERNAL), with a
message naming the field and the callable, since the fault is in host code
rather than the query. ast.analyze() called directly raises the analyzer's
check as a TypeError.
Timezone handling
DateRange bounds inside the AST are always timezone-aware UTC datetimes.
The tantivy emitter converts them to naive UTC before calling
Query.range_query (see _to_naive_utc in
src/whoosh_compat/emitters/tantivy_.py), because tantivy-py <= 0.26.0's
range_query only accepts naive datetimes for FieldType.Date. Passing a
tz-aware one raises ValueError. This is worked around here rather than
relied on upstream because the fix
(tantivy-py#666)
merged after 0.26.0 was tagged. Naive input is passed straight through
(tantivy already treats naive datetimes as UTC at index time, matching how
documents are indexed).
The JSON subpath carve-out
notes.user:alice-style dotted-field queries (FieldKind.JSON) are the one
place this library emits through index.parse_query() instead of
constructing a tantivy.Query programmatically: the installed tantivy-py's
Query.term_query cannot address a JSON subpath by exact field name: it
raises as if the field didn't exist at all. The emitter feature-detects this
per process and falls back to a strictly escaped/quoted parse_query call
for just that one leaf. This carve-out retires itself automatically once
tantivy-py#716 (which
routes make_term through JSON path resolution) lands and ships. No code
change is needed here, the feature-detection just starts taking the other
branch.
Development
Install with dev dependencies (uses uv):
uv sync --group dev
Run the checks CI runs:
uv run ruff check .
uv run mypy src
uv run pytest tests
This repo also ships a .pre-commit-config.yaml
covering the cheap, mechanical checks (whitespace, YAML/TOML validity,
codespell, zizmor for the GitHub Actions workflow, ruff check, keeping
uv.lock in sync). Run it with prek (a
faster, dependency-free reimplementation of pre-commit that reads the
same config file) or pre-commit itself:
uvx prek run --all-files
# or: uvx pre-commit run --all-files
Testing layers
- Unit tests (
tests/): parser, AST,normalize(), andFieldSpec/FieldRegistrybehavior in isolation. - Differential tests (
tests/differential/): parse the same corpus of query strings through both whoosh-compat and a real, pinned Whoosh (a test-only dependency; see the git-ref pin inpyproject.toml'sdevgroup, this fork carries parser fixes absent from the PyPI release) and compare the resulting trees. Divergences must match an explicit allowlist (tests/differential/allowlist.py), each entry cross- referencing aDIVERGENCES.mdentry for why it's expected. - End-to-end acceptance tests (
tests/emitter/test_acceptance_e2e.py): the same fixture documents indexed twice (once in a real Whoosh index, once in tantivy), full query strings run against both, and the matched document ID sets compared: result-level parity, not tree shape.
Because layer 2 needs a real Whoosh installation as an oracle, it's pulled in as a dev-only dependency (pinned by git ref, not the PyPI 2.7.4 release) rather than a runtime dependency of the library itself.
Property-based / fuzz testing
tests/differential/strategies.py is a grammar-aware Hypothesis
strategy covering the whole supported query language (README's syntax
table above): nested groups, every operator, wildcards/ranges/phrases/
comma-lists/boosts/JSON subpaths, and deliberately placed zero-token
values (an all-stopword term/phrase, see strategies.ZERO_TOKEN_WORDS).
It drives five properties:
tests/differential/test_hypothesis.py::test_fuzz_grammar_matches_oracle: the same AST-shape parity check as the static corpus (layer 2 above), but over generated, nested queries, guided byhypothesis.target()toward structurally rich examples (more nodes, deeper nesting, more distinct node types, more zero-token leaves buried inside a larger structure). Seeded with every static corpus line plus a few strings pulled directly fromDIVERGENCES.mdentries, via@example().tests/differential/test_hypothesis.py::test_normalize_is_total_and_idempotent:normalize(normalize(x)) == normalize(x)for every freshly parsed AST, andnormalize()never raises.tests/test_parse_never_raises.py::test_parse_raises_nothing_but_query_parser_error: over the same generated grammar (drawn deeper than the shared strategy's committed default),parse()raises nothing butQueryParserError, the documented "library defect" type. Sits alongside regression anchors for every escape route found so far.tests/emitter/test_hypothesis_e2e.py::test_emit_never_raises_except_unsupported: parsing a query that produced no diagnostics, then emitting it against a real in-memory tantivy index, never raises aQueryErrorwhosediagnostic.causeis anything other thanCause.UNSUPPORTED(the documented case of a construct that parses cleanly but has no way to execute against tantivy, such asDIVERGENCES.mdentry 5).tests/emitter/test_hypothesis_e2e.py::test_normalize_idempotent_on_emitter_registry_grammar: the samenormalize()property again, against the emitter registry's own (smaller, JSON/BOOLEAN_EXISTS-carrying) field vocabulary.
These run at a modest max_examples in CI so the suite stays fast. To run
a longer local soak (recommended before a release, or after touching
parser/, ast.py, or emitters/), raise the example count for a single
run without editing the files.
Note that a Hypothesis profile cannot do this. Registering a profile
with a higher max_examples and loading it has no effect here, because
every property in this repository sets max_examples explicitly in its
own @settings(...), and an explicit value beats the profile default no
matter when the profile is loaded. A run set up that way silently executes
the committed CI counts while appearing to run thousands of examples
(--hypothesis-show-statistics reports the real number, and says
Stopped because settings.max_examples=300).
What does work is rewriting the settings object the @given wrapper reads
at call time. Save this as soak_plugin.py somewhere on PYTHONPATH and
pass it with -p:
import os
import hypothesis
_TARGET = int(os.environ.get("SOAK_MAX_EXAMPLES", "5000"))
def pytest_collection_modifyitems(session, config, items):
for item in items:
fn = getattr(item, "obj", None)
current = getattr(fn, "_hypothesis_internal_use_settings", None)
if current is None or current.max_examples >= _TARGET:
continue
fn._hypothesis_internal_use_settings = hypothesis.settings(
current, max_examples=_TARGET, deadline=None
)
SOAK_MAX_EXAMPLES=5000 uv run pytest tests/differential/test_hypothesis.py tests/emitter/test_hypothesis_e2e.py -p soak_plugin -q --hypothesis-show-statistics
Two properties want their own (smaller) target rather than this one.
tests/emitter/test_acceptance_property.py's generated-query property
runs a real whoosh search and a real tantivy search per example, and
takes its count from the WHOOSH_COMPAT_ACCEPTANCE_SOAK_EXAMPLES
environment variable instead. And test_alternating_nesting_depth_cost_budget
asserts on elapsed wall-clock time, so it is marked wall_clock: any
runner that adds instrumentation must deselect it with
-m "not wall_clock".
Editing the max_examples values in those files directly, for the
duration of the run, also works and needs no plugin. Keep long soaks
(thousands of examples) out of CI: they take minutes, not seconds, and
are meant for local/pre-release verification, not every push.
HypoFuzz (coverage-guided fuzzing that runs existing Hypothesis
properties under instrumentation) was evaluated for this purpose and is
not used: its license (Zac-HD/hypofuzz's LICENSE, checked directly)
grants use "for non-commercial purposes only", explicitly requires a
separate paid commercial license for "use within a commercial
organization, including internal tooling or testing" and "use in
continuous integration or development pipelines for commercial products",
and prohibits modification/redistribution without permission. That's
incompatible with this BSD-2-Clause project (whoosh-compat is itself used
by, and expected to be run in CI by, commercial downstream users like
paperless-ngx installs), so it was not added as a dependency. The plain-
Hypothesis soak profile above is the recommended way to get a similar
"run longer, look harder" effect without it.
License
BSD-2-Clause, see LICENSE. This project's parser
(whoosh_compat/parser/) is forked from
whoosh-community/whoosh (a
fork of Matt Chaput's original Whoosh,
also on PyPI; the whoosh-community fork
is itself unmaintained). Forked files retain their original BSD-2-Clause
header. See NOTICE.
Metadata
Release files for whoosh-compat 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| whoosh_compat-0.3.0.tar.gz | 665.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| whoosh_compat-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 836.5 kB
Release files / whoosh_compat-0.3.0.tar.gz
| Download URL | whoosh_compat-0.3.0.tar.gz |
|---|---|
| Size | 665.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a7590ea0f178c60ca83774752d0d511c5d72b20357bbf361223aaab12e3d1813
|
|
BLAKE2b-256 checksum How to use checksums |
236b945311159351411dd5a1de82f2948fe1ef6bf36308ee65ed76f568b94a7a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.
Transparency logRelease files / whoosh_compat-0.3.0-py3-none-any.whl
| Download URL | whoosh_compat-0.3.0-py3-none-any.whl |
|---|---|
| Size | 170.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
91e22de08fc22be50fc90f836c72d45de9c7a552f61465ffd1ebee98a6b28cfb
|
|
BLAKE2b-256 checksum How to use checksums |
53b0538c7bdcd294ae32f1f6a13ab6486887e208567355820e4a24a56c4558c2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.
Transparency log