Skip to main content

VectorAmp Python SDK

Python client for the VectorAmp API. It is package-ready, typed, and defaults to https://api.vectoramp.com.

Licensed under the Apache License 2.0.

Install

pip install vectoramp

For local development:

git clone https://github.com/VectorAmp/vectoramp-python.git
cd vectoramp-python
pip install -e '.[dev]'
pytest

Authentication

The SDK sends API keys with the X-API-Key header.

from vectoramp import VectorAmp

client = VectorAmp(api_key="va_...")

Or use the environment variable:

export VECTORAMP_API_KEY=va_...
client = VectorAmp()

Configure a non-production API host with base_url:

client = VectorAmp(api_key="va_...", base_url="http://localhost:8080")

Datasets

Dataset creation always requests the SABLE index. The SDK intentionally does not expose an index_type option. For normal use, omit embedding config and VectorAmp uses the managed VectorAmp-Embedding-4B model with inferred dim 2560. Built-in helpers also infer dimensions for optional OpenAI BYOM models.

The only required argument is name. The default embedding is VectorAmp-Embedding-4B (dim=2560, metric="cosine"):

dataset = client.datasets.create("product-docs")

dataset_id = dataset.id  # also available as dataset["id"] for compatibility

Enable hybrid (dense + sparse) indexing with hybrid=True:

dataset = client.datasets.create("product-docs", hybrid=True)

Optional BYOM: use the openai helper only when you intentionally want an OpenAI embedding model (dimension is inferred):

from vectoramp import openai

dataset = client.datasets.create(
    "openai-docs",
    embedding=openai("small"),  # or openai("large")
)

To save or rotate the OpenAI API key as an organization secret first, then create the dataset with embedding.secret_ref, use the helper:

dataset = client.datasets.create_openai(
    name="product-docs",
    api_key="sk-...",
    secret_ref="emb:openai:api_key",  # default
)

# Or upsert the secret directly.
client.secrets.put_openai_api_key("sk-...")
client.org_secrets.update("emb:openai:api_key", "sk-rotated")

Declare typed metadata at creation, then merge fields or replace the complete schema:

from vectoramp import MetadataSchemaField

schema: list[MetadataSchemaField] = [
    {"name": "category", "type": "string"},
    {"name": "price", "type": "f32"},
]
dataset = client.datasets.create(name="products", metadata_schema=schema)
dataset.patch_metadata_schema([{"name": "inventory", "type": "u32"}])
dataset.replace_metadata_schema(schema)  # [] removes all declared fields

Canonical field types are string, u32, i32, i64, f32, and f64.

For a custom/unknown model you must pass dim explicitly:

dataset = client.datasets.create(
    "product-docs",
    embedding={"provider": "acme", "model": "acme-embed-v1"},
    dim=1024,
)

Create/get/list return Dataset resource objects. They keep the raw API payload and carry the client/services needed for instance methods:

dataset = client.datasets.get(dataset_id)
print(dataset.id, dataset.raw_data)

List methods return the API pagination envelope with Dataset objects in the datasets field:

page = client.datasets.list(limit=50, offset=0)
for dataset in page["datasets"]:
    print(dataset.id, dataset.raw_data)
print(page["total"], page["limit"], page["offset"])

Service-style methods are still available for callers that prefer passing dataset_id explicitly. Get or delete a dataset:

client.datasets.get(dataset_id)
client.datasets.delete(dataset_id)

Insert vectors and texts

Insert raw vectors. A record id may be a string or an integer; integer ids are sent as JSON numbers, not coerced to strings:

dataset.insert(
    [
        {
            "id": "doc-001",
            "values": [0.1, 0.2, 0.3],
            "metadata": {"title": "Intro", "source": "manual"},
        },
        {
            "id": 42,  # preserved as a JSON number
            "values": [0.4, 0.5, 0.6],
            "metadata": {"title": "Appendix"},
        },
    ],
)

Embed text with the dataset's configured model and insert the resulting vectors:

dataset.add_texts("VectorAmp uses SABLE for high-performance vector search.")

# Optional IDs and metadata are still supported for batches.
dataset.add_texts(
    ["VectorAmp uses SABLE for high-performance vector search."],
    ids=["sable-note"],
    metadatas=[{"source": "readme"}],
)

Delete vectors by id:

dataset.delete_vectors(["doc-001", 42])
client.datasets.delete_vectors(dataset_id, ["doc-001", 42])

Search

Search by text:

results = dataset.search("How does SABLE work?", top_k=10, include_documents=True)

Search by vector:

results = dataset.search([0.1, 0.2, 0.3], top_k=5)

Hybrid and filtered search:

results = dataset.search(
    text="wireless headphones",
    top_k=10,
    filters={"category": "electronics"},
    advanced_filters=[{"field": "price", "op": "lt", "value": 100}],
    hybrid=True,
    sparse_query="wireless headphones",
    alpha=0.7,
    rerank={"enabled": True},  # vectoramp / VectorAmp-Rerank-v1
)

Ingestion

Start ingestion from an existing source:

job = dataset.ingest_source("source-uuid")

Or pass a typed source builder. The SDK creates the source, extracts its returned ID, and starts the ingestion job for the dataset:

from vectoramp import WebSource

job = dataset.ingest_source(
    WebSource(
        start_urls=["https://docs.example.com/"],
        max_depth=1,
    )
)

The same one-liner works for any source type, e.g. Confluence:

from vectoramp import ConfluenceSource

job = dataset.ingest_source(
    ConfluenceSource(
        base_url="https://acme.atlassian.net",
        username="user@example.com",
        api_token="…",
        spaces=["ENG"],
    )
)

List jobs with pagination:

jobs = client.ingestion.list_jobs(dataset_id=dataset_id, limit=50, offset=0)

Create sources with typed helpers. client.sources is an alias for the ingestion source APIs, so existing client.ingestion.create_source(...) code still works:

web = client.sources.create_web(
    start_urls=["https://docs.example.com/"],
    max_depth=1,
    include_assets=True,
    max_assets_per_page=5,
)

s3 = client.sources.create_s3(
    bucket="my-bucket",
    prefix="documents/",
    role_arn="arn:aws:iam::123456789012:role/vectoramp-ingestion",
)

gcs = client.sources.create_gcs(bucket="my-gcs-bucket", prefix="documents/")

jira = client.sources.create_jira(
    cloud_id="atlassian-cloud-id",
    project_keys=["ENG"],
    include_comments=True,  # default
)

confluence = client.sources.create_confluence(
    cloud_id="atlassian-cloud-id",  # or base_url="https://acme.atlassian.net"
    username="user@example.com",
    api_token="…",                  # auth_mode defaults to "basic"
    spaces=["ENG", "DOCS"],         # empty/omitted = all accessible spaces
    include_attachments=True,       # default False
)

gdrive = client.sources.create_google_drive(
    folder_ids=["drive-folder-id"],
    include_shared_drives=True,
)

github = client.sources.create_github(
    installation_id=42,                 # from the VectorAmp GitHub App install
    repositories=["acme/api", "acme/web"],
    ref_mode="active",                  # server default: recently pushed branches
    include_pull_requests=True,         # default
)

gitlab = client.sources.create_gitlab(
    projects=["mygroup/myproject"],     # and/or groups=["mygroup"]
    auth_mode="token",                  # defaults to "oauth"
    access_token="…",
    gitlab_url="https://gitlab.com",    # override for self-managed
)

upload_source = client.sources.create_file_upload()

GitHub sources authenticate through the VectorAmp GitHub App, so the SDK never takes a GitHub token. Install the app from the Sources page in the VectorAmp app, then pass the installation_id it reports along with the owner/repo names you want ingested. Both connectors ingest repository files plus active branches, pull/merge requests, and review discussions; turn parts off with include_pull_requests / include_merge_requests, include_review_threads, and include_direct_commits, and narrow file selection with include_globs, exclude_globs, and max_file_size_bytes.

sync_mode is omitted unless you set it, so the server applies its default of "incremental" for the connectors that support it. Pass sync_mode="full" to force a full re-sync.

The supported typed source classes are WebSource, S3Source, GCSSource, GoogleDriveSource (source_type="gdrive"), JiraSource, ConfluenceSource, GitHubSource, GitLabSource, and FileUploadSource (source_type="file_upload"). Use GenericSource as an escape hatch when the API supports a source type before the SDK has a dedicated class:

from vectoramp import GenericSource

source = client.sources.create(
    GenericSource(
        name="custom-source",
        source_type="custom",
        config={"any_api_field": "value"},
    )
)

The low-level create API is preserved:

source = client.ingestion.create_source(
    name="docs-site",
    source_type="web",
    config={"start_urls": ["https://docs.example.com/"], "max_depth": 1},
)

Upload local files through the REST upload flow. The SDK creates a file_upload source with a generated name when source_name is omitted, initializes presigned uploads, uploads bytes to the returned URLs, and completes the upload job:

job = dataset.ingest_files(["./docs/whitepaper.pdf", "./docs/overview.txt"])

Pass source_name="product-docs-upload" only when you want a specific source name.

Intelligence / RAG

Non-streaming query:

answer = dataset.ask("What are the key product features?", top_k=5)
print(answer["answer"])

Scope a question to one or more datasets with dataset_ids. Omit it to search every dataset the API key can see:

answer = client.ask("Which contracts renew in Q4?", dataset_ids=[contracts.id, invoices.id])
everything = client.ask("What changed this week?")   # no dataset_ids -> all visible datasets

Streaming SSE query:

for event in client.ask_stream("Summarize the docs", dataset_ids=[dataset_id]):
    if event["chunk_type"] == "text":
        print(event["content"], end="")

Transport abstraction

VectorAmp depends on a small transport interface. RestTransport is provided today; a future gRPC transport can implement the same request, stream, and close methods without changing resource UX.

Development

pip install -e '.[dev]'
ruff check .
mypy src
pytest

CI runs Ruff, mypy, and pytest with coverage on every change.

Dataset documents

Datasets expose retained source documents when ingestion stored originals:

page = client.datasets.list_documents("dataset_id", limit=50, cursor=None, status="ready")
for document in page.get("documents", []):
    print(document["id"], document.get("file_name"))

content = client.datasets.download_document("dataset_id", "document_id")
open("document.bin", "wb").write(content)

# Resource-style calls work too:
dataset = client.datasets.get("dataset_id")
dataset.list_documents()
dataset.download_document("document_id")

Intelligence sessions

session = client.intelligence.create_session(title="Planning", dataset_id=dataset.id)
client.intelligence.append_message(session["id"], role="user", content="Summarize the docs")
messages = client.intelligence.list_messages(session["id"], limit=100)

Intelligence answers return sources[] and chunks[]. Inline [1] citations refer to sources[0]; preview_ref is an opaque preview token, not a storage key.

Method reference

Both access styles work everywhere the SDK allows it: client.datasets.search(id, …) and dataset.search(…). Required arguments are listed first; optional arguments show their default.

Client (VectorAmp)

  • VectorAmp(api_key=None, *, base_url="https://api.vectoramp.com", timeout=30.0)api_key falls back to VECTORAMP_API_KEY.
  • client.ask(query, *, dataset_ids=None, top_k=5, conversation_history=None, include_sources=True)dataset_ids is a list; omit it to search every visible dataset.
  • client.ask_stream(query, *, dataset_ids=None, top_k=5, conversation_history=None, include_sources=True) — iterator of SSE chunks.
  • client.close() (also a context manager).

Datasets (client.datasets / Dataset)

  • create(name, *, dim=None, metric="cosine", embedding=None, embedding_provider="vectoramp", embedding_model="VectorAmp-Embedding-4B", hybrid=False, filters=None, metadata_schema=None, tuning=None, openai_api_key=None, openai_secret_ref="emb:openai:api_key", validate_openai_key=False)Dataset. Always SABLE. dim inferred for built-in models; required for custom models.
  • list(*, limit=50, offset=0) → page with Dataset objects.
  • get(dataset_id)Dataset.
  • patch_metadata_schema(dataset_id, schema) / Dataset.patch_metadata_schema(schema) → updated Dataset.
  • replace_metadata_schema(dataset_id, schema) / Dataset.replace_metadata_schema(schema) → updated Dataset.
  • delete(dataset_id) / dataset.delete().
  • stats(dataset_id) / dataset.stats().
  • search(dataset_id, query=None, *, vector=None, text=None, search_text=None, top_k=10, filters=None, advanced_filters=None, embedding_provider=None, embedding_model=None, nprobe_override=None, rerank_depth_override=None, hybrid=None, sparse_query=None, alpha=None, include_embeddings=None, include_documents=None, include_metadata=None, rerank=None) / dataset.search(…). query accepts a string (text) or float sequence (vector); top_k defaults to 10.
  • insert(dataset_id, vectors) and insert_vectors(dataset_id, vectors) / dataset.insert(vectors). Record id may be str or int (integers stay JSON numbers).
  • delete_vectors(dataset_id, ids, *, write_concern=None) / dataset.delete_vectors(ids).
  • embed(dataset_id, *, text=None, texts=None) / dataset.embed(…).
  • add_texts(dataset_id, texts, *, ids=None, metadatas=None) / dataset.add_texts(texts, …). Single string or list; ids auto-generated when omitted; copies the text into metadata.text.
  • list_documents(dataset_id, *, limit=50, cursor=None, status=None) / dataset.list_documents(…).
  • download_document(dataset_id, document_id) / dataset.download_document(document_id) → bytes.
  • ensure_engine(dataset_id).
  • dataset.ask(query, *, top_k=5, conversation_history=None, include_sources=True).
  • dataset.ingest_source(source, *, pipeline_id=None)source is an id or a source builder.
  • dataset.ingest_files(paths, *, source_name=None, description=None).

Sources / ingestion (client.sources is an alias of client.ingestion)

  • create(source) / create_source(source=None, *, name=None, source_type=None, config=None, description=None, metadata=None).
  • create_web(*, start_urls, name=None, max_depth=None, max_pages=None, allowed_domains=None, include_patterns=None, exclude_patterns=None, crawl_delay_seconds=None, include_assets=None, max_assets_per_page=None, sync_mode omitted via builder, description=None, metadata=None, config_extra=None).
  • create_s3(*, bucket, name=None, region="us-east-1", prefix=None, sync_mode=None, access_key_id=None, secret_access_key=None, role_arn=None, endpoint_url=None, file_patterns=None, max_file_size_mb=None, …).
  • create_gcs(*, bucket, name=None, prefix=None, project_id=None, credentials_json=None, sync_mode=None, file_patterns=None, max_file_size_mb=None, …).
  • create_jira(*, cloud_id, name=None, access_token=None, project_keys=None, jql=None, include_comments=True, sync_mode=None, …).
  • create_confluence(*, cloud_id=None, base_url=None, name=None, auth_mode="basic", username=None, api_token=None, oauth_credentials=None, spaces=None, include_attachments=False, sync_mode=None, …) — requires cloud_id or base_url.
  • create_google_drive(*, name=None, folder_ids=None, file_ids=None, auth_mode="oauth", oauth_credentials=None, include_shared_drives=None, sync_mode=None, service_account_json=None, credentials_json=None, …).
  • create_file_upload(*, name="vectoramp-python-upload", storage_provider="s3", sync_mode="full", …).
  • list_sources(*, limit=50, offset=0), get_source(source_id).
  • start_job(*, source_id, dataset_id, pipeline_id=None), list_jobs(*, dataset_id=None, limit=50, offset=0), get_job(job_id), retry_job(job_id).
  • ingest_files(*, dataset_id, paths, source_name=None, description=None), init_upload(source_id, files), complete_upload(source_id, *, job_id, file_ids).

Builder classes: WebSource, S3Source, GCSSource, GoogleDriveSource, JiraSource, ConfluenceSource, FileUploadSource, GenericSource (escape hatch). sync_mode is omitted unless set, so the server default ("incremental") applies.

Schedules (client.schedules)

  • list(*, limit=50, offset=0), get(schedule_id).
  • create(*, source_id, dataset_id, cron, timezone=None, pipeline_id=None, enabled=None, name=None, metadata=None).
  • update(schedule_id, *, cron=None, timezone=None, pipeline_id=None, enabled=None, name=None, metadata=None) — only passed fields change.
  • delete(schedule_id), trigger(schedule_id).

Intelligence (client.intelligence)

  • query(query, *, dataset_id=None, top_k=5, conversation_history=None, include_sources=True).
  • stream(query, *, …) — iterator of SSE chunks.
  • create_session(*, title=None, workspace_id=None, dataset_id=None, metadata=None).
  • list_sessions(*, limit=50), get_session(session_id).
  • append_message(session_id, *, role, content, metadata=None), list_messages(session_id, *, limit=100).

License

Apache License 2.0. See LICENSE and NOTICE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vectoramp-0.5.1.tar.gz (48.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vectoramp-0.5.1-py3-none-any.whl (38.2 kB view details)

Uploaded Python 3

File details

Details for the file vectoramp-0.5.1.tar.gz.

File metadata

  • Download URL: vectoramp-0.5.1.tar.gz
  • Upload date:
  • Size: 48.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vectoramp-0.5.1.tar.gz
Algorithm Hash digest
SHA256 8983c8cbbeada25ba53c12b22b1433a7e8bd90b50cf1e5ef4c199d9c28caaa00
MD5 fb9da76947c4af0992f73da515e3d756
BLAKE2b-256 944b5809c7ffdfd883780b2bf85a928fece370c70e479bfbf30659d3e5243bd8

See more details on using hashes here.

Provenance

The following attestation bundles were made for vectoramp-0.5.1.tar.gz:

Publisher: publish.yml on VectorAmp/vectoramp-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vectoramp-0.5.1-py3-none-any.whl.

File metadata

  • Download URL: vectoramp-0.5.1-py3-none-any.whl
  • Upload date:
  • Size: 38.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vectoramp-0.5.1-py3-none-any.whl
Algorithm Hash digest
SHA256 f66efd18450502b5c579151a512b0adae5d9d3e6d265fdf72e94ef111aedaf11
MD5 240f35025b1beba1104cadfc572f19e6
BLAKE2b-256 f4c9fc1c4b4ee507526212ccd96902a463bcebc15c010d090a3c170913925c3a

See more details on using hashes here.

Provenance

The following attestation bundles were made for vectoramp-0.5.1-py3-none-any.whl:

Publisher: publish.yml on VectorAmp/vectoramp-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.5.1 This release

2 files

0.5.0

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page