vaultspec-rag
The semantic search component for vault and code.
Search code and feature records by meaning through the command line or Model Context Protocol (MCP). Search inference runs on your GPU, with optional hosted Typesafe query classification and result reranking.
Install · Use it · Docs · Help
Use it with vaultspec-core or independently in another repository. To index PDFs and other formats, connect a converter.
What you need
For the Python installation below, use Python 3.13 or 3.14 and uv. Only the process that hosts the inference service needs model packages and an accelerator. It requires NVIDIA CUDA on Linux or Windows, or Apple silicon on macOS; CPU inference and AMD GPUs are unsupported. A command-line or MCP client that connects to an already-running service on the same machine does not need CUDA.
Check the memory and disk requirements before installing. That section also covers the smaller resource profile.
Install
Choose extras for what you want this environment to run. There is no rag extra.
| Role | Package | Loads models here? | Needs an accelerator? |
|---|---|---|---|
| Command-line client and service controls | vaultspec-rag |
No | No |
| MCP stdio adapter to an existing service | vaultspec-rag[mcp] |
No | No |
| Inference-service host | vaultspec-rag[gpu] |
Yes | Yes |
| Inference host with local MCP adapter | vaultspec-rag[gpu,mcp] |
Yes | Yes |
The client and MCP adapter use the compatible vaultspec-rag HTTP service listening on
the configured loopback port. A remote Qdrant URL moves vector storage only; it is not
a remote inference service and does not remove the host's gpu requirement. See the
installation lanes for setup
commands and the limits of each role.
Install a standalone tool for use across repositories. Choose the command for your platform. The commands below install both the inference service and MCP adapter. These CUDA commands use Python 3.13 and pin the GPU wheel so later tool upgrades retain it.
Windows x64:
uv tool install --python 3.13 "vaultspec-rag[gpu,mcp]" --with "torch @ https://download.pytorch.org/whl/cu130/torch-2.14.0%2Bcu130-cp313-cp313-win_amd64.whl"
Linux x86_64 (glibc 2.28 or newer):
uv tool install --python 3.13 "vaultspec-rag[gpu,mcp]" --with "torch @ https://download.pytorch.org/whl/cu130/torch-2.14.0%2Bcu130-cp313-cp313-manylinux_2_28_x86_64.whl"
Apple silicon macOS:
uv tool install --python 3.13 "vaultspec-rag[gpu,mcp]"
For other Python versions or Linux architectures, see GPU wheel selection.
If uv reports that its executables directory is missing from PATH, follow its
instructions before continuing. For an existing tool installation, follow the
upgrade instructions before
replacing its environment.
Once installation succeeds, open the repository you want to search.
The default setup downloads
naver/splade-v3, a gated sparse model.
Before running it, accept the model's access conditions and authenticate the service
account with HF_TOKEN or hf auth login; a token alone is insufficient until its
account has accepted the conditions. The model's CC-BY-NC-SA-4.0 license restricts
commercial use. If the gate or license is unsuitable, follow the
dense-only setup instead.
vaultspec-rag install --no-torch-config
This installs the repository's agent integration, downloads the three search models,
and provisions Qdrant, the index server. The GPU packages are already installed, so
--no-torch-config leaves the project's PyTorch configuration alone. The first setup
downloads several gigabytes; subsequent projects share the models and server binary.
The default installer sets up the local models and Qdrant even with the lightweight
base or [mcp] package. Use install --no-provision to connect a client-only workspace
to an already-running service.
Check the installation:
vaultspec-rag server doctor
Check that the report detects your GPU and finds all three models and the Qdrant binary. If it reports a problem, use the installation troubleshooting guide.
Use it
Optional Typesafe classification
Set VAULTSPEC_RAG_TYPESAFE_API_KEY in the service account's environment before
starting the server to opt into paid Typesafe classification. A valid, funded key
enables query interpretation and reranking using the full result content, including
removal of confidently irrelevant hits. The server sends queries and candidate
content to Typesafe; without a usable key, search keeps its existing local ranking.
server start and server status show the running server's enrollment and whether
a recent evaluation succeeded. No separate enable flag is needed. See
activation, fallback and status meanings
before enabling it. The local search models and GPU are still required.
Start and search
Start the service to load the models. The command waits until it is ready:
vaultspec-rag server start
Index the repository from its root:
vaultspec-rag index
Wait for indexing to finish before searching. Use vaultspec-rag server jobs --watch
to follow progress. The service watches for file changes and updates the index
automatically afterwards.
Search source code with --type code, or feature records with --type vault:
vaultspec-rag search "parse query text into filters" --type code
If results are missing or incomplete, check the index and adjust the query.
One service handles all your repositories. Run index in each repository you want
to search. See index maintenance for rebuilding or
removing indexed content.
Other ways to install
To share a version with collaborators, add RAG as a project dependency.
Without a Python toolchain, use the prebuilt Windows or Linux binaries.
Remove RAG
Follow the removal guide to preview project changes, choose whether to clean up indexes, and remove the package.
Refine searches
- Choose query terms.
- Filter by language, path, or document type.
- Investigate missing results and index coverage.
- Monitor indexing jobs.
Use it from an AI assistant
Follow MCP setup to connect your coding agent. The default toolset includes tools that change or delete indexes. To restrict access, see withholding the mutating tools.
Use an on-disk index
By default, vaultspec-rag uses managed Qdrant in a separate process. The optional local-only backend keeps an embedded on-disk index inside the RAG process. It still requires a GPU and models.
Switching backends does not migrate your existing index. Follow backend setup.
Read PDFs and other formats
Converters extract content from unsupported formats for indexing. See converter setup.
Converters run without a sandbox, with the permissions of the account running RAG. They can access files and the network. They can run during explicit indexing, watched changes, and agent-triggered reindexing.
Before indexing, inspect .vaultragpreprocess.toml and its commands. Use
vaultspec-rag preprocess status to inspect configuration without running converters.
See security and disable options.
Scripting it
For JSON output and result handling, follow scripting and automation.
Documentation
- Run your first search
- Installation and troubleshooting
- Worked searches
- Commands and flags
- Configuration reference
- Architecture and indexing internals
- Release notes
Status and help
vaultspec-rag is Beta. Report issues with your version, operating system, GPU, command, and error output. Redact credentials and private content before posting.
Related projects
- vaultspec-core: Decision-driven harness for coding agents, and humans.
- vaultspec-dashboard: The human-facing visual workspace for a Vaultspec project.
License
vaultspec-rag is released under the MIT License.
Metadata
Release files for vaultspec-rag 0.4.35
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vaultspec_rag-0.4.35.tar.gz | 6.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vaultspec_rag-0.4.35-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 7.5 MB
Release files / vaultspec_rag-0.4.35.tar.gz
| Download URL | vaultspec_rag-0.4.35.tar.gz |
|---|---|
| Size | 6.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1c21f0c236e868f1551c3c7a9313828794599d04167469b03a1c3989e2197488
|
|
BLAKE2b-256 checksum How to use checksums |
befe333a1d39aa8e1832b0308b93361c23528f72803b682d2c72ce9777d0fc8f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.17 {"installer":{"name":"uv","version":"0.12.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / vaultspec_rag-0.4.35-py3-none-any.whl
| Download URL | vaultspec_rag-0.4.35-py3-none-any.whl |
|---|---|
| Size | 1.5 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3ecd600809302804a4997a2ba49f1437166189cf33c6a216c1a4dfdcb93dc299
|
|
BLAKE2b-256 checksum How to use checksums |
5611e36828a9eabd8ef77a578b41852b63391ea00a6e1fd4dd5e021914606c36
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.17 {"installer":{"name":"uv","version":"0.12.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|