Skip to main content

Sifts – Simple Full Text & Semantic Search

🔎 Sifts is a simple but powerful Python package for managing and querying document collections with support for both SQLite and PostgreSQL databases.

It is designed to efficiently handle full-text search and vector search, making it ideal for applications that involve large-scale text data retrieval.

Features

  • Dual Database Support: Sifts works with both SQLite and PostgreSQL, offering the simplicity of SQLite for lightweight applications and the scalability of PostgreSQL for larger, production environments.
  • Full-Text Search (FTS): Perform advanced text search queries with full-text search support.
  • Vector Search: Integrate with embedding models to perform vector-based similarity searches, perfect for applications involving natural language processing.
  • Flexible Querying: Supports complex queries with filtering, ordering, and pagination.

Background

The main idea of Sifts is to leverage the built-in full-text search capabilities in SQLite and PostgreSQL and to make them available via a unified, Pythonic API. You can use SQLite for small projects or development and trivially switch to PostgreSQL to scale your application.

For vector search, cosine similarity is computed in PostgreSQL via the pgvector extension, while with SQLite similarity is calculated in memory.

Sifts does not come with a server mode as it's meant as a library to be imported by other apps. The original motivation for its development was to replace whoosh as search backend in Gramps Web, which is based on Flask.

Installation

You can install Sifts via pip:

pip install sifts

Usage

import sifts

# by default, creates a new SQLite database in the working directory
collection = sifts.Collection(name="my_collection")

# Add docs to the index. Can also update and delete.
collection.add(
    documents=["Lorem ipsum dolor", "sit amet"],
    metadatas=[{"foo": "bar"}, {"foo": "baz"}], # otpional, can filter on these
    ids=["doc1", "doc2"], # unique for each doc. Uses UUIDs if omitted
)

results = collection.query(
    "Lorem",
    # limit=2,  # optionally limit the number of results
    # where={"foo": "bar"},  # optional filter
    # order_by="foo",  # sort by metadata key (rather than rank)
)

The API is inspired by chroma.

Full-text search syntax

Sifts supports the following search syntax:

  • Search for individual words
  • Search for multiple words (will match documents where all words are present)
  • and operator
  • or operator
  • * wildcard (in SQLite, supported anywhere in the search term, in PostgreSQL only at the end of the search term)

The search syntax is the same regardless of backend.

Special Character Handling

Sifts automatically handles special characters that have meaning in SQLite FTS5 and PostgreSQL tsquery syntax, preventing syntax errors when searching for terms containing these characters.

Automatically Quoted Characters (SQLite)

The following characters are automatically quoted when found in search terms:

  • Parentheses () - used for grouping in FTS5
  • Brackets [] - used for column filters in FTS5
  • Curly braces {} - used for advanced FTS5 syntax
  • Colon : - used for column-specific searches
  • Comma , - used as token separator in FTS5
  • Quote " - used for phrase searches
  • Hyphen - - in hyphenated words like "test-word"
  • Apostrophe ' - in contractions like "it's"

Examples

# Search for city with comma - automatically quoted
collection.query("Bydgoszcz, Poland")
# Internal query: "Bydgoszcz," Poland

# Search for time with colon - automatically quoted
collection.query("time:12:00")
# Internal query: "time:12:00"

# Search with parentheses - automatically quoted
collection.query("test (example)")
# Internal query: test "(example)"

# Wildcards work with special characters
collection.query("city:*")
# Internal query: "city:"*

Manual Quoting

You can still use explicit quotes for exact phrase matching:

# Exact phrase match
collection.query('"exact phrase"')

To include a literal quote in your search, use FTS5's double-quote escape:

# Search for: test"value
collection.query('"test""value"')

PostgreSQL Differences

PostgreSQL tsquery handles special characters differently than SQLite FTS5:

  • SQLite: Quotes special characters to avoid FTS5 syntax errors
  • PostgreSQL: Replaces special characters with spaces to preserve term boundaries

Example:

# SQLite
collection.query("time:12:00")
# Internal: "time:12:00" (quoted to avoid syntax error)

# PostgreSQL
collection.query("time:12:00")
# Internal: time & 12 & 00 (split into separate words)

Both approaches work because full-text search tokenization typically strips punctuation during indexing, so documents containing these characters can still be found by searching for the words they contain.

Sifts can also be used as vector store, used for semantic search engines or retrieval-augmented generation (RAG) with large language models (LLMs).

Simply pass the embedding_function to the Collection factory to enable vector storage and set vector_search=True in the query method. For instance, using the Sentence Transformers library,

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("intfloat/multilingual-e5-small")

def embedding_function(queries: list[str]):
    return model.encode(queries)

collection = sifts.Collection(
    db_url="sqlite:///vector_store.db",
    name="my_vector_store",
    embedding_function=embedding_function
)

# Adding vector data to the collection
collection.add(["This is a test sentence.", "Another example query."])

# Querying the collection with semantic search
results = collection.query("Find similar sentences.", vector_search=True)

Some embedding models expect queries and documents to be embedded differently (e.g. with different prefixes). In this case, pass a separate query_embedding_function, which is used for query strings, while embedding_function is used for documents:

collection = sifts.Collection(
    db_url="sqlite:///vector_store.db",
    name="my_vector_store",
    embedding_function=model.encode_document,
    query_embedding_function=model.encode_query,
)

If you have already computed the document vectors, you can pass them to add or update with the embeddings argument instead of having them computed by embedding_function:

collection.add(["This is a test sentence."], embeddings=[vector])

PostgreSQL collections require installing and enabling the pgvector extension.

Updating and Deleting Documents

Documents can be updated or deleted using their IDs.

# Update a document
collection.update(ids=["document_id"], contents=["Updated content"])

# Delete a document
collection.delete(ids=["document_id"])

Contributing

Contributions are welcome! Feel free to create an issue if you encounter problems or have an improvement suggestion, and even better submit a PR along with it!

License

Sifts is licensed under the MIT License. See the LICENSE file for details.


Happy Sifting! 🚀

Metadata

Release files for sifts 1.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sifts 1.4.0
File Size Uploaded
sifts-1.4.0.tar.gz 28.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sifts 1.4.0
File Interpreter ABI Platform
sifts-1.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 41.7 kB

Release files / sifts-1.4.0.tar.gz

Download URL sifts-1.4.0.tar.gz
Size 28.2 kB
Tags Source
SHA-256 checksum
How to use checksums
014966a2f9aa0c02ff94f2fdf46cae845b6e3beb827b25b55741e7502504746b
BLAKE2b-256 checksum
How to use checksums
c04e3c0a66a1fd7f69a1ea38e514faa0562708eeff0cfe3b1c24be51ab4ed959
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / sifts-1.4.0-py3-none-any.whl

Download URL sifts-1.4.0-py3-none-any.whl
Size 13.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a681bf114e345484b71ebef8082ef169b1c5ef29eba5d3006ca4594cebb18db1
BLAKE2b-256 checksum
How to use checksums
57d9f04b3bb965f12ad2635727dbda47edceae00b7b91183e0ae488bcd0ed577
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

1.4.1

2 release files

This release

1.4.0 This release

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5

2 release files

0.4

2 release files

0.3

2 release files

0.2

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page