Skip to main content

Moss client library for Python

moss enables private, on-device semantic search in your Python applications with cloud storage capabilities.

Built for developers who want instant, memory-efficient, privacy-first AI features with seamless cloud integration.

✨ Features

  • ⚡ On-Device Vector Search - Sub-millisecond retrieval with zero network latency
  • 🔍 Semantic, Keyword & Hybrid Search - Embedding search blended with Keyword matching
  • ☁️ Cloud Storage Integration - Automatic index synchronization with cloud storage
  • 📦 Multi-Index Support - Manage multiple isolated search spaces
  • 🛡️ Privacy-First by Design - Computation happens locally, only indexes sync to cloud
  • 🚀 High-Performance Rust Core - Built on optimized Rust bindings for maximum speed
  • 🧠 Custom Embedding Overrides - Provide your own document and query vectors when you need full control

📦 Installation

pip install moss

🚀 Quick Start

import asyncio
from moss import MossClient, DocumentInfo, QueryOptions

async def main():
    # Initialize search client with project credentials
    client = MossClient("your-project-id", "your-project-key")

    # Prepare documents to index
    documents = [
        DocumentInfo(
            id="doc1",
            text="How do I track my order? You can track your order by logging into your account.",
            metadata={"category": "shipping"}
        ),
        DocumentInfo(
            id="doc2", 
            text="What is your return policy? We offer a 30-day return policy for most items.",
            metadata={"category": "returns"}
        ),
        DocumentInfo(
            id="doc3",
            text="How can I change my shipping address? Contact our customer service team.",
            metadata={"category": "support"}
        )
    ]

    # Create an index with documents (syncs to cloud)
    index_name = "faqs"
    await client.create_index(index_name, documents)  # Defaults to moss-minilm
    print("Index created and synced to cloud!")

    # Load the index (from cloud or local cache)
    await client.load_index(index_name)

    # Search the index
    result = await client.query(
        index_name,
        "How do I return a damaged product?",
        QueryOptions(top_k=3, alpha=0.6),
    )

    # Display results
    print(f"Query: {result.query}")
    for doc in result.docs:
        print(f"Score: {doc.score:.4f}")
        print(f"ID: {doc.id}")
        print(f"Text: {doc.text}")
        print("---")

asyncio.run(main())

Scores and min_score

Each result's score is its relevance to the query from 0 to 1. Results are ordered by hybrid relevance, so a lower result can show a higher score. Scores depend on the embedding model: tune min_score on your own queries.

QueryOptions(min_score=0.5) drops results whose score is below 0.5; a query returns up to top_k results that clear it, in hybrid order. It takes a value from 0 to 1. Keyword-only queries (alpha=0.0) have no embedding to score against: every result scores 0.0 and min_score raises an error.

Candidate depth

A hybrid query ranks a fixed number of hits on each signal (vector and keyword) and fuses those two lists, so a document outside both lists is never returned. By default it looks 2 * top_k deep. QueryOptions(candidate_depth=200) looks deeper, which raises recall and query time. It takes a value of at least 1 and is raised to top_k when below it. Keyword-only and vector-only queries (alpha 0 or 1) rank one signal, where a deeper list returns the same results, so they ignore it. SessionIndex.query ignores it too.

🔥 Example Use Cases

  • Smart knowledge base search with cloud backup
  • Realtime Voice AI agents with persistent indexes
  • Personal note-taking search with sync across devices
  • Private in-app AI features with cloud storage
  • Local semantic search in edge devices, fully on-device

Available Models

  • moss-minilm: Lightweight model optimized for speed and efficiency
  • moss-mediumlm: Balanced model offering higher accuracy with reasonable performance

🔧 Getting Started

Prerequisites

  • Python 3.8 or higher
  • Valid InferEdge project credentials

Environment Setup

  1. Install the package:
pip install moss
  1. Get your credentials:

Sign up at InferEdge Platform to get your project_id and project_key.

  1. Set up environment variables (optional):
export MOSS_PROJECT_ID="your-project-id"
export MOSS_PROJECT_KEY="your-project-key"
# Optional: override the manage API host (defaults to https://service.usemoss.dev)
export MOSS_CLOUD_API_BASE_URL="https://service.usemoss.dev"

Basic Usage

import asyncio
from moss import MossClient, DocumentInfo, QueryOptions

async def main():
    # Initialize client
    client = MossClient("your-project-id", "your-project-key")
    
    # Create and populate an index
    documents = [
        DocumentInfo(id="1", text="Python is a programming language"),
        DocumentInfo(id="2", text="Machine learning with Python is popular"),
    ]
    
    await client.create_index("my-docs", documents)
    await client.load_index("my-docs")
    
    # Search
    results = await client.query(
        "my-docs",
        "programming language",
        QueryOptions(alpha=1.0),
    )
    for doc in results.docs:
        print(f"{doc.id}: {doc.text} (score: {doc.score:.3f})")

asyncio.run(main())

Hybrid Search Controls

alpha lets you decide how much weight to give semantic similarity versus keyword relevance when running query():

# Pure keyword search
await client.query("my-docs", "programming language", QueryOptions(alpha=0.0))

# Mixed results (default 0.8 => semantic heavy)
await client.query("my-docs", "programming language")

# Pure embedding search
await client.query("my-docs", "programming language", QueryOptions(alpha=1.0))

Pick any value between 0.0 and 1.0 to tune the blend for your use case.

Disk cache

Pass cache_path to persist the downloaded index to disk. Later loads reuse the cached copy while the cloud version is unchanged, so restarts skip the download. Auto-refresh writes through to the same cache. Each load still contacts the cloud to check for a newer version.

await client.load_index(
    "my-docs",
    auto_refresh=True,
    polling_interval_in_seconds=300,
    cache_path="/var/cache/moss",
)

Set cache_path once on the client to make it the default for every load. A per-call cache_path overrides it, and both override the native default (~/.moss). The chosen directory also holds the .moss-device-id file that keys Monthly Active Device billing.

client = MossClient(
    "your-project-id",
    "your-project-key",
    cache_path="/var/cache/moss",
)

# Uses /var/cache/moss.
await client.load_index("my-docs", auto_refresh=True)

# Overrides it for this load only.
await client.load_index("other-docs", cache_path="/tmp/moss")

Offline-tolerant load

With a cache_path set and a snapshot already on disk, a load whose cloud metadata check fails on a transport error (unreachable network, timeout) serves the cached snapshot instead of failing, so a restart during an outage still answers queries. The probe uses a few-second budget, not the full retry window, and the load starts a refresh poller so the index self-heals when the cloud returns. A real "index not found" (404) still fails, so a deleted index does not resurrect from cache.

A load served this way opens the index, but a built-in embedding model cannot be opened until the Moss cloud is reachable, so text queries with alpha above 0 fail until then. Keyword-only queries (alpha=0.0) and queries that pass their own embedding keep working, and a later refresh or query warms the model once the cloud returns.

was_served_stale(name) reports whether an index is currently served from such a snapshot; it clears once a refresh reconfirms it against the cloud.

await client.load_index("my-docs", cache_path="/var/cache/moss")
if await client.was_served_stale("my-docs"):
    # Serving a cached copy; the cloud was unreachable at load.
    ...

Search several loaded indexes in one call and get the global top-K back, with each result tagged by its source index_name. All indexes must be loaded locally and share the same embedding model.

loaded = await client.load_indexes(["products", "reviews"])
if not loaded.loaded:
    raise RuntimeError(f"no indexes loaded: {loaded.failed}")

results = await client.query_multi_index(
    loaded.loaded,
    "noise cancelling headphones",
    QueryOptions(top_k=5, alpha=0.5),
)
for doc in results.docs:
    print(f"[{doc.index_name}] {doc.id}: {doc.text} (score: {doc.score:.3f})")

await client.unload_indexes(loaded.loaded)

alpha works exactly as in query() (default 0.8): 1.0 is embedding-only, 0.0 is keyword-only, and anything in between blends both with Reciprocal Rank Fusion. Keyword scoring runs each index's own BM25 and merges each index's raw hits before fusion; because BM25 statistics stay per-corpus, cross-index keyword ranking is approximate. top_k caps the merged result, not each index, and filter applies to every index. load_indexes is best-effort: names that fail are reported in failed without rolling back the ones that loaded.

Web sources

POST the existing /v1/manage web-source actions and return typed results. Poll job_id with get_job_status. Depends on the index-manager release with several sources per index and source-preserving crawls, and the moss-control release with the deleteIndex proxy. Set MOSS_CLOUD_API_BASE_URL to send manage calls to a non-default host (defaults to https://service.usemoss.dev).

created = await client.create_web_source(
    "https://docs.example.com",
    "knowledge-base",
    max_pages=200,
    refresh_cadence="weekly",
)
await client.get_job_status(created.job_id)

sources = await client.list_web_sources(index_name="knowledge-base")
source = await client.get_web_source(created.id)
updated = await client.update_web_source(created.id, refresh_cadence="daily")
# updated.job_id is set when resync=True enqueued a crawl.
resync = await client.resync_web_source(created.id)
deleted = await client.delete_web_source(created.id)
# deleted.purge_job_id is set when a purge was enqueued.

Metadata filtering

You can pass a metadata filter directly to query() after loading an index locally:

results = await client.query(
    "my-docs",
    "running shoes",
    QueryOptions(top_k=5, alpha=0.6),
    filter={
        "$and": [
            {"field": "category", "condition": {"$eq": "shoes"}},
            {"field": "price", "condition": {"$lt": "100"}},
        ]
    },
)

For a complete runnable example, see python/user-facing-sdk/samples/metadata_filtering.py.

🧠 Providing custom embeddings

Already using your own embedding model? Supply vectors directly when managing indexes and queries:

import asyncio

from moss import DocumentInfo, MossClient, QueryOptions


def my_embedding_model(text: str) -> list[float]:
    """Placeholder for your custom embedding generator."""
    ...


async def main() -> None:
    client = MossClient("your-project-id", "your-project-key")

    documents = [
        DocumentInfo(
            id="doc-1",
            text="Attach a caller-provided embedding.",
            embedding=my_embedding_model("doc-1"),
        ),
        DocumentInfo(
            id="doc-2",
            text="Fallback to the built-in model when the field is omitted.",
            embedding=my_embedding_model("doc-2"),
        ),
    ]

    await client.create_index("custom-embeddings", documents)  # Defaults to moss-minilm
    await client.load_index("custom-embeddings")

    results = await client.query(
        "custom-embeddings",
        "<query text>",
        QueryOptions(embedding=my_embedding_model("<query text>"), top_k=10),
    )

    print(results.docs[0].id, results.docs[0].score)


asyncio.run(main())

Leaving the model argument undefined defaults to moss-minilm. Pass QueryOptions to reuse your own embeddings or to override top_k on a per-query basis.

Telemetry

The SDK reports usage to the Moss cloud at $MOSS_INDEX_URL/telemetry (https://service.usemoss.dev/index/telemetry by default). Telemetry is always on, because the device id and the query count drive usage billing. Requests are sent in the background. A failed send is dropped and never fails the call that caused it.

Queries do not send a request each. The SDK buffers their events, with the same fields as below, and sends them every 3 seconds while queries run, plus once more when the client or session is released. A send puts up to 500 query events in one JSON array per request. The buffer holds up to 10,000 events between sends, and events past that are dropped. Other events are sent one per request, when the operation happens.

Fields on every event

Field Value
action Always "telemetry".
eventType The event type, from the table below.
projectId Your project id.
clientId, sessionId The same random id, one per client or session object.
sdkVersion, sdkPackageVersion Version of the native binding package.
nativeCoreVersion Version of the Rust core.
modelCacheSchemaVersion Layout version of the on-disk model cache.
deviceId, mossDeviceId The Moss device UUID stored in .moss-device-id under cache_path or ~/.moss. deviceId is the billable one.
indexName The index, when the event concerns one index.
modelId, modelArtifactVersion, modelManifestSha256 The embedding model and its exact artifact, when known.

Event types and their extra fields

Event Sent when Extra fields
index.load An index loads docCount, autoRefresh, refreshIntervalSecs (with auto refresh), cached (a cache path was given), servedStale (served from the cache while the cloud is unreachable), modelCacheGeneration
index.auto_refresh An auto refresh poll finishes changed, modelCacheGeneration
index.unload An index unloads or the client closes modelCacheGeneration
index.query A single-index query runs, sent in a batch latencyMs, modelCacheGeneration
index.query_multi A multi-index query runs, sent in a batch The index.query fields, plus indexCount and indexNames. modelCacheGeneration when every index shares one, else modelCacheGenerationsByIndex.
index.usage.periodic_flush, index.usage.final_flush Every 3 seconds while there is usage, and at close queryCount, docsEmbedded, tokensEstimated
session.load A session loads an index from the cloud docCount
session.load_from_disk A session loads a local snapshot docCount, plus rebuilt and persisted when a legacy snapshot is rebuilt
session.add_docs, session.delete_docs Documents are added or deleted docCount
session.push_index A session pushes its index docCount
session.auto_refresh_staged A newer cloud version is downloaded updatedAt
session.auto_refresh A downloaded version is installed changed, docCount, updatedAt
session.unload A session is released None
session.query A session query runs, sent in a batch latencyMs
session.usage.periodic_flush, session.usage.final_flush Every 3 seconds while there is usage, at release, and before a push queryCount, docsEmbedded, tokensEstimated

latencyMs on a query event is the time that query took, in milliseconds. index.query and index.query_multi time the whole call, including the query embedding. session.query times the search lookup only, without embedding, so the two are not directly comparable.

No document text, query text, metadata or embedding is ever sent.

📄 License

This package is licensed under the PolyForm Shield License 1.0.0.

  • ✅ Free for any use, including production and commercial use.
  • ❌ Not permitted: providing a product that competes with Moss.
  • 📩 For commercial licenses, contact: contact@usemoss.dev

📬 Contact

For support, commercial licensing, or partnership inquiries, contact us: contact@usemoss.dev

Metadata

Release files for moss 1.14.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for moss 1.14.0
File Size Uploaded
moss-1.14.0.tar.gz 34.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for moss 1.14.0
File Interpreter ABI Platform
moss-1.14.0-py3-none-any.whl Python 3 none any Details

Total release size: 64.6 kB

Release files / moss-1.14.0.tar.gz

Download URL moss-1.14.0.tar.gz
Size 34.1 kB
Tags Source
SHA-256 checksum
How to use checksums
273176b02b1d1dd9731e0b2ff9a2b36ae073c5d6a6473d2c43bb96947862940f
BLAKE2b-256 checksum
How to use checksums
e44ab6c6fd5c9050089271cdc0c51a7d144a75fd010ee878cdddb8c575408d49
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / moss-1.14.0-py3-none-any.whl

Download URL moss-1.14.0-py3-none-any.whl
Size 30.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
75417623434e46ecca4c773a47c76d7eebf52eb68bc217ad176164716c286030
BLAKE2b-256 checksum
How to use checksums
6fcf88fbfe9d0391bb62bb8118110e41febde442ab87917f0d8df5111c09e282
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

This release

1.14.0 This release

2 release files

1.13.0

2 release files

1.12.0

2 release files

1.11.0

2 release files

1.10.0

2 release files

1.9.1

2 release files

1.9.0

2 release files

1.8.0

2 release files

1.7.3

2 release files

1.7.2

2 release files

1.7.1

2 release files

1.7.0

2 release files

1.6.0

2 release files

1.5.0

2 release files

1.4.1

2 release files

1.4.0

2 release files

1.2.0

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page