txtai

Build AI-powered semantic search applications

These details have not been verified by PyPI

Project links

Project description

Build AI-powered semantic search applications

txtai executes machine-learning workflows to transform data and build AI-powered semantic search applications.

demo

Traditional search systems use keywords to find data. Semantic search applications have an understanding of natural language and identify results that have the same meaning, not necessarily the same keywords.

Backed by state-of-the-art machine learning models, data is transformed into vector representations for search (also known as embeddings). Innovation is happening at a rapid pace, models can understand concepts in documents, audio, images and more.

Summary of txtai features:

🔎 Large-scale similarity search with multiple index backends (Faiss, Annoy, Hnswlib)
📄 Create embeddings for text snippets, documents, audio, images and video. Supports transformers and word vectors.
💡 Machine-learning pipelines to run extractive question-answering, zero-shot labeling, transcription, translation, summarization and text extraction
↪️️ Workflows that join pipelines together to aggregate business logic. txtai processes can be microservices or full-fledged indexing workflows.
⚙️ Build with Python or YAML. API bindings available for JavaScript, Java, Rust and Go.
☁️ Cloud-native architecture that scales out with container orchestration systems (e.g. Kubernetes)

Applications range from similarity search to complex NLP-driven data extractions to generate structured databases. Semantic workflows transform and find data driven by user intent.

flows

The following applications are powered by txtai.

apps

Application	Description
paperai	AI-powered literature discovery and review engine for medical/scientific papers
tldrstory	AI-powered understanding of headlines and story text
neuspo	Fact-driven, real-time sports event and news site
codequestion	Ask coding questions directly from the terminal

txtai is built with Python 3.7+, Hugging Face Transformers, Sentence Transformers and FastAPI

Why txtai?

why

In addition to traditional search systems, a growing number of semantic search solutions are available, so why txtai?

Up and running in minutes with pip or Docker

# Get started in a couple lines
from txtai.embeddings import Embeddings

embeddings = Embeddings({"path": "sentence-transformers/all-MiniLM-L6-v2"})
embeddings.index([(0, "Correct", None), (1, "Not what we hoped", None)])
embeddings.search("positive", 1)
#[(0, 0.2986203730106354)]

Build applications in your programming language of choice via the API

# app.yml
embeddings:
    path: sentence-transformers/all-MiniLM-L6-v2

CONFIG=app.yml uvicorn "txtai.api:app"
curl -X GET "http://localhost:8000/search?query=positive"

Connect machine learning models together to build intelligent data processing workflows
Works with both small and big data - scale when needed
Low footprint - install additional dependencies when you need them
Learn by example - notebooks cover all available functionality

Installation

install

The easiest way to install is via pip and PyPI

pip install txtai

Python 3.7+ is supported. Using a Python virtual environment is recommended.

See the detailed install instructions for more information covering optional dependencies, environment specific prerequisites, installing from source and how to run with containers.

Examples

examples

The examples directory has a series of notebooks and applications giving an overview of txtai. See the sections below.

Semantic Search

Build semantic/similarity/vector/neural search applications.

Notebook	Description
Introducing txtai ▶️	Overview of the functionality provided by txtai
Build an Embeddings index with Hugging Face Datasets	Index and search Hugging Face Datasets
Build an Embeddings index from a data source	Index and search a data source with word embeddings
Add semantic search to Elasticsearch	Add semantic search to existing search systems
Similarity search with images	Embed images and text into the same space for search
Distributed embeddings cluster	Distribute an embeddings index across multiple data nodes
What's new in txtai 4.0	Content storage, SQL, object storage, reindex and compressed indexes
Anatomy of a txtai index	Deep dive into the file formats behind a txtai embeddings index
Custom Embeddings SQL functions	Add user-defined functions to Embeddings SQL
Model explainability	Explainability for semantic search
Query translation	Domain-specific natural language queries with query translation
Build a QA database	Question matching with semantic search
Embeddings components	Composable search with vector, SQL and scoring components
Semantic Graphs	Explore topics, data connectivity and run network analysis

Pipelines

Transform data with NLP-backed pipelines.

Notebook	Description
Extractive QA with txtai	Introduction to extractive question-answering with txtai
Extractive QA with Elasticsearch	Run extractive question-answering queries with Elasticsearch
Extractive QA to build structured data	Build structured datasets using extractive question-answering
Apply labels with zero shot classification	Use zero shot learning for labeling, classification and topic modeling
Building abstractive text summaries	Run abstractive text summarization
Extract text from documents	Extract text from PDF, Office, HTML and more
Transcribe audio to text	Convert audio files to text
Translate text between languages	Streamline machine translation and language detection
Generate image captions and detect objects	Captions and object detection for images
Near duplicate image detection	Identify duplicate and near-duplicate images
API Gallery	Using txtai in JavaScript, Java, Rust and Go

Workflows

Efficiently process data at scale.

Notebook	Description
Run pipeline workflows ▶️	Simple yet powerful constructs to efficiently process data
Transform tabular data with composable workflows	Transform, index and search tabular data
Tensor workflows	Performant processing of large tensor arrays
Entity extraction workflows	Identify entity/label combinations
Workflow Scheduling	Schedule workflows with cron expressions
Push notifications with workflows	Generate and push notifications with workflows
Pictures are a worth a thousand words	Generate webpage summary images with DALL-E mini
Run txtai with native code	Execute workflows in native code with the Python C API

Model Training

Train NLP models.

Notebook	Description
Train a text labeler	Build text sequence classification models
Train without labels	Use zero-shot classifiers to train new models
Train a QA model	Build and fine-tune question-answering models
Export and run models with ONNX	Export models with ONNX, run natively in JavaScript, Java and Rust
Export and run other machine learning models	Export and run models from scikit-learn, PyTorch and more

Applications

Series of example applications with txtai. Links to hosted versions on Hugging Face Spaces also provided.

Application	Description
Basic similarity search	Basic similarity search example. Data from the original txtai demo.	🤗
Book search	Book similarity search application. Index book descriptions and query using natural language statements.	Local run only
Image search	Image similarity search application. Index a directory of images and run searches to identify images similar to the input query.	🤗
Summarize an article	Summarize an article. Workflow that extracts text from a webpage and builds a summary.	🤗
Wiki search	Wikipedia search application. Queries Wikipedia API and summarizes the top result.	🤗
Workflow builder	Build and execute txtai workflows. Connect summarization, text extraction, transcription, translation and similarity search pipelines together to run unified workflows.	🤗

Documentation

Full documentation on txtai including configuration settings for pipelines, workflows, indexing and the API.

Contributing

For those who would like to contribute to txtai, please see this guide.

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

9.6.0

Feb 25, 2026

9.5.0

Feb 12, 2026

9.4.1

Jan 23, 2026

9.4.0

Jan 21, 2026

9.3.0

Dec 23, 2025

9.2.0

Nov 21, 2025

9.1.0

Nov 4, 2025

9.0.1

Sep 15, 2025

9.0.0

Aug 28, 2025

8.6.0

Jun 10, 2025

8.5.0

Apr 14, 2025

8.4.0

Mar 11, 2025

8.3.1

Feb 12, 2025

8.3.0

Feb 11, 2025

8.2.0

Jan 9, 2025

8.1.0

Dec 10, 2024

8.0.0

Nov 18, 2024

7.5.1

Oct 25, 2024

7.5.0

Oct 14, 2024

7.4.0

Sep 5, 2024

7.3.0

Jul 15, 2024

7.2.0

May 31, 2024

7.1.0

Apr 19, 2024

7.0.0

Feb 21, 2024

6.3.0

Jan 2, 2024

6.2.0

Nov 8, 2023

6.1.0

Sep 26, 2023

6.0.0

Aug 10, 2023

5.5.1

Apr 27, 2023

5.5.0

Apr 20, 2023

5.4.0

Mar 6, 2023

5.3.0

Feb 7, 2023

5.2.0

Dec 20, 2022

5.1.0

Oct 18, 2022

This version

5.0.0

Sep 27, 2022

4.6.0

Aug 15, 2022

4.5.0

May 17, 2022

4.4.0

Apr 20, 2022

4.3.1

Mar 11, 2022

4.3.0

Mar 10, 2022

4.2.1

Feb 28, 2022

4.2.0

Feb 24, 2022

4.1.0

Feb 3, 2022

4.0.0

Jan 11, 2022

3.7.0

Nov 23, 2021

3.6.0

Nov 8, 2021

3.5.0

Oct 18, 2021

3.4.0

Oct 7, 2021

3.3.0

Sep 10, 2021

3.2.0

Aug 17, 2021

3.1.0

May 22, 2021

3.0.0

May 4, 2021

2.0.0

Jan 13, 2021

1.5.0

Nov 21, 2020

1.4.0

Nov 3, 2020

1.3.0

Oct 11, 2020

1.2.1

Sep 11, 2020

1.2.0

Sep 10, 2020

1.1.0

Aug 18, 2020

1.0.0

Aug 11, 2020

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

txtai-5.0.0.tar.gz (101.5 kB view details)

Uploaded Sep 27, 2022 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

txtai-5.0.0-py3-none-any.whl (147.4 kB view details)

Uploaded Sep 27, 2022 Python 3

File details

Details for the file txtai-5.0.0.tar.gz.

File metadata

Download URL: txtai-5.0.0.tar.gz
Upload date: Sep 27, 2022
Size: 101.5 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.2.0 pkginfo/1.5.0.1 requests/2.26.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.62.3 CPython/3.7.14

File hashes

Hashes for txtai-5.0.0.tar.gz
Algorithm	Hash digest
SHA256	`e7506fa0d97e5a5dca5155dd9afc647d06828af4daefb42105429421f472d012`
MD5	`c7288eeaffb87eca7d777244adff9c5b`
BLAKE2b-256	`651ebf3fb18a50ffd74408061b8e7be227502bb95316003b331169a1912041ae`

See more details on using hashes here.

File details

Details for the file txtai-5.0.0-py3-none-any.whl.

File metadata

Download URL: txtai-5.0.0-py3-none-any.whl
Upload date: Sep 27, 2022
Size: 147.4 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.2.0 pkginfo/1.5.0.1 requests/2.26.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.62.3 CPython/3.7.14

File hashes

Hashes for txtai-5.0.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`c5a598e8575fc8299d6332862354e489a2f28c84472a9530a2d35611adbfa1af`
MD5	`ecb1717998b2241752c929b7a20f82e4`
BLAKE2b-256	`12f5b50b99ff76f800eeb22b82f5d5d7818ca28127ccabc3d0bea5f767d9d686`

See more details on using hashes here.

txtai 5.0.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Build AI-powered semantic search applications

Why txtai?

Installation

Examples

Semantic Search

Pipelines

Workflows

Model Training

Applications

Documentation

Further Reading

Contributing

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes