Skip to main content
vaultspec-rag logo

vaultspec-rag

The semantic search component for vault and code.

Grep finds a concept only when you already know its name. Why code looks the way it does is often written in a decision record that's hard to find. vaultspec-rag searches a repository's source code and its decision records by meaning. Run it from the command line, or connect an AI assistant through the Model Context Protocol (MCP), so the assistant finds both the code and the decisions behind it. Search combines a model that matches meaning with one that matches exact terms, then reranks the results. An optional paid service, Typesafe, can refine the ranking.

One background service per machine runs the models on the local graphics processing unit (GPU), because they're too slow to be useful on a central processing unit (CPU). A host installation starts that service; a client installation only sends it requests. The architecture overview explains how the pieces fit.

GitHub Stars GitHub Forks Watchers Branches Contributors Last commit Commits Open issues Closed issues Open PRs Closed PRs Merged PRs Release Build Runtime License Agent-friendly AGENTS.md

Install · Use it · Docs · Help

Use it with vaultspec-core or independently in another repository. To index PDFs and other formats, connect a converter.

What you need

  • Python 3.13 or 3.14 with uv, or a prebuilt binary.
  • A supported GPU: an NVIDIA GPU with CUDA, NVIDIA's GPU computing platform, on Linux or Windows, or Apple silicon on macOS. vaultspec-rag doesn't run on a CPU or on AMD GPUs.
  • Enough memory for a resource profile. The default profile needs 16 GiB of system memory and, on CUDA, 12 GiB of free GPU memory. The smaller embedded-local profile needs 8 GiB of system memory and 6 GiB of free GPU memory.
  • Several gigabytes of disk for a one-time model download.
  • By default, search also uses the sparse model naver/splade-v3, which matches exact terms. It needs a Hugging Face account that has accepted the model's non-commercial licence. The dense-only setup uses only the meaning model. It needs no Hugging Face login and gives up exact-term matching.

The installation requirements list the full figures, and the glossary defines the terms used here.

Install

  • If vaultspec-rag isn't installed on this machine yet, install the host. It serves every repository on the machine and gives AI assistants the search tools.
  • If a host installation already runs the service, add a client to any uv-managed Python project that must list vaultspec-rag as a dependency. A client installs no GPU packages or models and sends every request to the host's service on the same machine.

The installation guide covers every route, and its troubleshooting, upgrade, and removal sections cover what comes after.

Host installation

Install the host once, as a standalone tool; it serves every repository on the machine. The [gpu] extra adds PyTorch and the model libraries, and [mcp] adds the MCP adapter that AI assistants launch. The Windows and Linux commands record the CUDA package index in the tool's installation receipt, so uv tool upgrade keeps resolving the GPU build.

Windows x64 and Linux x86_64 or aarch64 (glibc 2.28 or newer):

uv tool install --python 3.13 "vaultspec-rag[gpu,mcp]" --index https://download.pytorch.org/whl/cu130 --index-strategy unsafe-first-match

Apple silicon macOS, which uses Metal rather than CUDA:

uv tool install --python 3.13 "vaultspec-rag[gpu,mcp]"

Upgrade later with uv tool upgrade vaultspec-rag, then restart the service so it runs the new release; if uv reports an entry point it could not overwrite because the file is in use, the release is installed and the running launcher keeps working. An installation made without the two index options resolves a CPU-only PyTorch at its next upgrade; vaultspec-rag server doctor reports that and prints the two commands that repair it in place, which pin the GPU build describes. If uv reports that its executables directory isn't on your PATH, run uv tool update-shell and open a new terminal.

By default, search needs access to the sparse model. If you can't accept its licence, set VAULTSPEC_RAG_SPARSE_ENABLED=0 in your user environment and skip to the repository setup. Otherwise, accept the licence on the model page, then log in:

uvx --from huggingface_hub hf auth login

Alternatively, set HF_TOKEN in your user environment.

From the root of each repository you want to search, run the repository setup. It doesn't reinstall the tool:

vaultspec-rag install --no-torch-config

The setup adds the AI-assistant integration and creates the .vault/ folder for decision records. On the first repository, it also downloads the search models and Qdrant, the index server; later repositories reuse both. The first run downloads several gigabytes. --no-torch-config leaves the repository's own PyTorch configuration alone, because the tool already carries its PyTorch.

Start the service. It loads the models and waits until it's ready:

vaultspec-rag server start

The service doesn't start by itself after a reboot, so run vaultspec-rag server start again then. To stop it, run vaultspec-rag server stop.

Check the installation:

vaultspec-rag server doctor

vaultspec-rag server doctor - service, GPU, model, and Qdrant readiness at a glance

Check that the report detects your GPU and finds every configured model and the Qdrant binary. If it reports a problem, use the installation troubleshooting guide.

Client in a Python project

A client lets a project's own AI-assistant configuration launch the search tools from the project environment, so collaborators get them with uv sync. Each collaborator still needs their own host installation at the release the project pins.

  1. In the host installation, run this command and note the release on the Service release: line:

    vaultspec-rag server status --verbose
    
  2. From the project root, add the client pinned to that release:

    uv add --dev "vaultspec-rag[mcp]==<release>"
    
  3. Set up the project. --mode dev makes the AI assistant launch the search tools from the project environment, even if the host installation set up the project first. A client skips PyTorch configuration and all downloads.

    uv run vaultspec-rag install --mode dev
    
  4. Confirm the client reaches the service. uv run vaultspec-rag server doctor must show release: <release> (matches this client) and report PyTorch as not needed for this client installation. If the release doesn't match, repeat step 2 with the release from step 1.

Run the client as uv run vaultspec-rag, and start or stop the service from the host installation. When you upgrade the host, move each client to the same release with the upgrade steps.

Use it

From the root of each repository you want to search, index it. The command queues indexing jobs and prints their IDs:

vaultspec-rag index

Follow progress with vaultspec-rag server jobs --watch, and wait until the jobs finish before searching. Afterwards, the service watches for file changes and updates the index automatically.

Search source code with --type code, or decision records with --type vault. Results list file paths with their matching passages:

vaultspec-rag search "parse query text into filters" --type code

A client runs the same commands with the uv run prefix. If results are missing or incomplete, check the index and adjust the query. The getting-started tutorial walks through a first search, and AI assistant setup connects your AI assistant. See index maintenance for rebuilding or removing indexed content.

Optional Typesafe classification

Set VAULTSPEC_RAG_TYPESAFE_API_KEY in the service account's environment before starting the server to opt into paid Typesafe classification. A valid, funded key enables query interpretation and reranking using the full result content, including removal of confidently irrelevant hits. The server sends queries and candidate content to Typesafe; without a usable key, search keeps its existing local ranking.

server start and server status show the running server's enrollment and whether a recent evaluation succeeded. No separate enable flag is needed. See activation, fallback, and status meanings before enabling it. The local search models and GPU are still required.

Other ways to install

To have one Python project's environment run the service, install the host as a project dependency. Unlike a client, this adds the GPU packages to the project.

Without a Python toolchain, use the prebuilt Windows, Linux, or Apple silicon macOS binaries.

Remove vaultspec-rag

Follow the removal guide to preview project changes, choose whether to clean up indexes, and remove the package.

Refine searches

Use it from an AI assistant

Follow MCP setup to connect your coding agent. The default toolset includes tools that change or delete indexes. To restrict access, see withholding the mutating tools.

Use an on-disk index

By default, the service keeps its index in managed Qdrant, a separate process. The optional local-only backend keeps the index in each repository's .vault/ folder instead. It still needs a GPU and the models. To use it, set VAULTSPEC_RAG_INDEX_SUPPORT_PROFILE=embedded-local in the environment that starts the service, because the default resource profile refuses the local-only backend.

Switching backends doesn't migrate your existing index. Follow backend setup to switch.

Read PDFs and other formats

Converters extract content from unsupported formats for indexing. See converter setup.

Converters run without a sandbox, with the permissions of the account running RAG. They can access files and the network. They can run during explicit indexing, watched changes, and agent-triggered reindexing.

Before indexing, inspect .vaultragpreprocess.toml and its commands. Use vaultspec-rag preprocess status to inspect configuration without running converters. See security and disable options.

Scripting it

For JSON output and result handling, follow scripting and automation.

Documentation

Status and help

vaultspec-rag is Beta. Report issues with your version, operating system, GPU, command, and error output. Redact credentials and private content before posting.

License

vaultspec-rag is released under the MIT License.

Metadata

Release files for vaultspec-rag 0.5.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vaultspec-rag 0.5.3
File Size Uploaded
vaultspec_rag-0.5.3.tar.gz 6.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for vaultspec-rag 0.5.3
File Interpreter ABI Platform
vaultspec_rag-0.5.3-py3-none-any.whl Python 3 none any Details

Total release size: 7.9 MB

Release files / vaultspec_rag-0.5.3.tar.gz

Download URL vaultspec_rag-0.5.3.tar.gz
Size 6.3 MB
Tags Source
SHA-256 checksum
How to use checksums
c738ab7914f0d216ca41422618f203953941aea4e9980b395d0a69306a461359
BLAKE2b-256 checksum
How to use checksums
0b14a5bfcc91185814ffa1daa2d3ae647ddbcef9e9f086c133e6189a7b38e23f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / vaultspec_rag-0.5.3-py3-none-any.whl

Download URL vaultspec_rag-0.5.3-py3-none-any.whl
Size 1.6 MB
Tags Python 3
SHA-256 checksum
How to use checksums
e295b2fd718f906fac0d60898556938b92e0f32304d05c36247d259e7d8f4360
BLAKE2b-256 checksum
How to use checksums
e3126d1af925104e48a69ef52defec7d40ea6bd84ba2264427185ebfbb11098c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.6.0

2 release files

This release

0.5.3 This release

2 release files

0.4.35

2 release files

0.4.34

2 release files

0.4.21

2 release files

0.4.20

2 release files

0.4.19

2 release files

0.4.18

2 release files

0.4.17

2 release files

0.4.16

2 release files

0.4.15

2 release files

0.4.14

2 release files

0.4.13

2 release files

0.4.12

2 release files

0.4.11

2 release files

0.4.10

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.14

2 release files

0.3.13

2 release files

0.3.12

2 release files

0.3.11

2 release files

0.3.10

2 release files

0.3.9

2 release files

0.3.8

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.27

2 release files

0.2.26

2 release files

0.2.25

2 release files

0.2.24

2 release files

0.2.23

2 release files

0.2.22

2 release files

0.2.21

2 release files

0.2.20

2 release files

0.2.19

2 release files

0.2.18

2 release files

0.2.17

2 release files

0.2.10

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page