Python SDK for the Vector Gateway service (embeddings and vector search)
Project description
Vector SDK for Python
A lightweight Python client for submitting embedding requests and vector search queries to the Vector Gateway service.
Overview
The Vector SDK provides a simple interface for generating embeddings via the centralized Vector Gateway service. The SDK communicates directly with Redis Streams, making it efficient and suitable for any Python service that can reach the shared Redis VM.
Key Features:
- Simple, Pythonic API with namespace-based organization
- Intuitive methods:
client.embeddings,client.search,client.db - Asynchronous request submission with optional waiting
- Full type hints and documentation
- Multiple embedding model support (Google Vertex AI and OpenAI)
- Client-side model validation before submission
- Minimal dependencies (just Redis)
Installation
From Source (Monorepo)
cd packages/py/vector-sdk
pip install -e .
# Or with uv
uv pip install -e .
From Package Registry (when published)
pip install sf-vector-sdk
Authentication
All SDK operations require a valid API key. Contact your administrator to obtain an API key.
from vector_sdk import VectorClient
client = VectorClient(
redis_url="redis://your-redis-host:6379",
http_url="http://localhost:8080",
api_key="vsk_v1_your_api_key_here", # Required
)
API Key Format: vsk_v1_{32_random_chars}
Unauthenticated Usage: Some utility functions work without VectorClient or API key:
from vector_sdk import compute_content_hash, extract_tool_text
# These work offline - no API key required
hash_val = compute_content_hash(
"FlashCard",
{"type": "BASIC", "term": "ATP", "definition": "Adenosine triphosphate"}
)
Quick Start
Basic Usage
from vector_sdk import VectorClient
import os
# Create client with API key
client = VectorClient(
redis_url="redis://your-redis-host:6379",
http_url="http://localhost:8080",
api_key=os.environ["VECTOR_API_KEY"], # Required
)
# Create embeddings
result = client.embeddings.create_and_wait(
texts=[
{"id": "doc1", "text": "Introduction to machine learning"},
{"id": "doc2", "text": "Deep neural networks explained"},
],
content_type="topic",
)
print(f"Processed: {result.processed_count}, Failed: {result.failed_count}")
# Vector search
search_result = client.search.query_and_wait(
query_text="What is machine learning?",
database="turbopuffer",
namespace="topics",
top_k=10,
)
for match in search_result.matches:
print(f"{match.id}: {match.score}")
# Direct database lookup (no embedding)
docs = client.db.get_by_ids(
ids=["doc1"],
database="turbopuffer",
namespace="topics",
)
client.close()
With Storage Configuration
from vector_sdk import VectorClient, StorageConfig, MongoDBStorage, TurboPufferStorage
client = VectorClient(redis_url="redis://your-redis-host:6379")
# Create embeddings with storage configuration
result = client.embeddings.create_and_wait(
texts=[
{
"id": "tool123",
"text": "Term: Photosynthesis. Definition: The process by which plants convert sunlight into energy.",
"document": {
"toolId": "tool123",
"toolCollection": "FlashCard",
"userId": "user456",
"contentHash": "abc123",
}
}
],
content_type="flashcard",
priority="high",
storage=StorageConfig(
mongodb=MongoDBStorage(
database="events_new",
collection="tool_vectors",
embedding_field="toolEmbedding",
upsert_key="contentHash",
),
turbopuffer=TurboPufferStorage(
namespace="tool_vectors",
id_field="_id",
metadata=["toolId", "toolCollection", "userId"],
),
),
metadata={"source": "my-service"},
)
client.close()
Context Manager
with VectorClient(redis_url="redis://localhost:6379") as client:
result = client.embeddings.create_and_wait(
texts=[{"id": "doc1", "text": "Hello world"}],
content_type="document",
)
# Connection automatically closed
API Reference
VectorClient
The main client class providing namespaced access to all SDK functionality.
Constructor
client = VectorClient(
redis_url="redis://localhost:6379",
http_url="http://localhost:8080", # Optional, required for db operations
api_key="vsk_v1_your_api_key", # Required
)
Parameters:
redis_url(str, required): Redis connection URLhttp_url(str, optional): HTTP URL for db operationsapi_key(str, required): API key for authentication
Namespaces
client.embeddings
Embedding generation operations.
| Method | Description |
|---|---|
create(texts, content_type, ...) |
Submit embedding request, return request ID |
wait_for(request_id, timeout) |
Wait for request completion |
create_and_wait(texts, content_type, ...) |
Submit and wait for result |
get_queue_depth() |
Get current queue depth for each priority |
# Async: create and wait separately
request_id = client.embeddings.create(texts, content_type)
result = client.embeddings.wait_for(request_id)
# Sync: create and wait in one call
result = client.embeddings.create_and_wait(texts, content_type)
# Check queue depth
depths = client.embeddings.get_queue_depth()
client.search
Vector similarity search operations.
| Method | Description |
|---|---|
query(query_text, database, ...) |
Submit search query, return request ID |
wait_for(request_id, timeout) |
Wait for query completion |
query_and_wait(query_text, database, ...) |
Submit and wait for result |
# Vector search with semantic similarity
result = client.search.query_and_wait(
query_text="What is machine learning?",
database="turbopuffer",
namespace="topics",
top_k=10,
include_metadata=True,
)
client.db
Direct database operations (no embedding required). Requires http_url.
| Method | Description |
|---|---|
get_by_ids(ids, database, ...) |
Lookup documents by ID |
find_by_metadata(filters, database, ...) |
Search by metadata filters |
clone(id, source_namespace, destination_namespace) |
Clone document between namespaces |
delete(id, namespace) |
Delete document from namespace |
client.structured_embeddings
Type-safe embedding for known tool types (FlashCard, TestQuestion, etc.) with automatic text extraction, content hash computation, and database routing.
| Method | Description |
|---|---|
embed_flashcard(data, metadata) |
Embed a flashcard, return request ID |
embed_flashcard_and_wait(data, metadata, timeout) |
Embed and wait for result |
embed_flashcard_batch(items) |
Embed batch of flashcards, return request ID |
embed_flashcard_batch_and_wait(items, timeout) |
Embed batch and wait for result |
embed_test_question(data, metadata) |
Embed a test question, return request ID |
embed_test_question_and_wait(data, metadata, timeout) |
Embed and wait for result |
embed_test_question_batch(items) |
Embed batch of test questions, return request ID |
embed_test_question_batch_and_wait(items, timeout) |
Embed batch and wait for result |
embed_spaced_test_question(data, metadata) |
Embed a spaced test question, return request ID |
embed_spaced_test_question_and_wait(data, metadata, timeout) |
Embed and wait for result |
embed_spaced_test_question_batch(items) |
Embed batch of spaced test questions, return request ID |
embed_spaced_test_question_batch_and_wait(items, timeout) |
Embed batch and wait for result |
embed_audio_recap(data, metadata) |
Embed an audio recap section, return request ID |
embed_audio_recap_and_wait(data, metadata, timeout) |
Embed and wait for result |
embed_audio_recap_batch(items) |
Embed batch of audio recaps, return request ID |
embed_audio_recap_batch_and_wait(items, timeout) |
Embed batch and wait for result |
embed_topic(data, metadata) |
Embed a topic (uses TopicMetadata), return request ID |
embed_topic_and_wait(data, metadata, timeout) |
Embed and wait for result (uses TopicMetadata) |
embed_topic_batch(items) |
Embed batch of topics (uses TopicMetadata), return request ID |
embed_topic_batch_and_wait(items, timeout) |
Embed batch and wait for result (uses TopicMetadata) |
Metadata Types:
ToolMetadata- For tools (FlashCard, TestQuestion, etc.) - requirestool_idTopicMetadata- For topics only - all fields optional (user_id,topic_id)
from vector_sdk import VectorClient, ToolMetadata, TopicMetadata, TestQuestionInput
client = VectorClient(redis_url="redis://localhost:6379")
# Embed a flashcard - uses ToolMetadata (tool_id required)
result = client.structured_embeddings.embed_flashcard_and_wait(
data={"type": "BASIC", "term": "Mitochondria", "definition": "The powerhouse of the cell"},
metadata=ToolMetadata(tool_id="tool123", user_id="user456", topic_id="topic789"),
)
# Embed a test question - uses ToolMetadata (tool_id required)
result = client.structured_embeddings.embed_test_question_and_wait(
data=TestQuestionInput(
question="What is the capital?",
answers=[...],
question_type="multiplechoice",
),
metadata=ToolMetadata(tool_id="tool456"),
)
# Embed a topic - uses TopicMetadata (all fields optional)
# Note: Topic data requires an "id" field which becomes the TurboPuffer document ID
result = client.structured_embeddings.embed_topic_and_wait(
data={"id": "topic-123", "topic": "Photosynthesis", "description": "The process by which plants convert sunlight to energy"},
metadata=TopicMetadata(user_id="user123", topic_id="topic456"), # No tool_id needed
)
# Batch embedding - embed multiple topics in a single request
from vector_sdk import TopicBatchItem
batch_result = client.structured_embeddings.embed_topic_batch_and_wait(
items=[
TopicBatchItem(data={"id": "topic-1", "topic": "Topic 1", "description": "Description 1"}, metadata=TopicMetadata(user_id="user1")),
TopicBatchItem(data={"id": "topic-2", "topic": "Topic 2", "description": "Description 2"}, metadata=TopicMetadata(topic_id="topic2")),
TopicBatchItem(data={"id": "topic-3", "topic": "Topic 3", "description": "Description 3"}, metadata=TopicMetadata()), # All optional
],
)
Database Routing:
Set the STRUCTURED_EMBEDDING_DATABASE_ROUTER environment variable:
| Value | Behavior |
|---|---|
dual |
Write to both TurboPuffer AND Pinecone if both have enabled: True |
turbopuffer |
Only write to TurboPuffer |
pinecone |
Only write to Pinecone |
| undefined | Defaults to turbopuffer |
# Lookup by IDs
result = client.db.get_by_ids(
ids=["doc1", "doc2"],
database="turbopuffer",
namespace="topics",
)
# Find by metadata
result = client.db.find_by_metadata(
filters={"userId": "user123"},
database="mongodb",
collection="vectors",
database_name="mydb",
)
# Clone between namespaces
result = client.db.clone("doc1", "ns1", "ns2")
# Delete
result = client.db.delete("doc1", "ns1")
# Export entire namespace
export_result = client.db.get_vectors_in_namespace(
namespace="tool_vectors",
include_vectors=True,
)
print(f"Exported {len(export_result.documents)} documents")
Types
Result Types
@dataclass
class EmbeddingResult:
request_id: str
status: str # "success", "partial", "failed"
processed_count: int
failed_count: int
errors: list[EmbeddingError]
timing: Optional[TimingBreakdown]
completed_at: datetime
@property
def is_success(self) -> bool: ...
@property
def is_partial(self) -> bool: ...
@property
def is_failed(self) -> bool: ...
@dataclass
class QueryResult:
request_id: str
status: str # "success", "failed"
matches: list[VectorMatch]
error: Optional[str]
timing: Optional[QueryTiming]
completed_at: datetime
@dataclass
class VectorMatch:
id: str
score: float # Similarity score (0-1, higher is more similar)
metadata: Optional[dict]
vector: Optional[list[float]]
Priority Levels
| Priority | Use Case | Description |
|---|---|---|
critical |
Real-time user requests | Reserved quota, processed first |
high |
New content embeddings | Standard processing priority |
normal |
Updates, re-embeddings | Default priority |
low |
Backfill, batch jobs | Processed when capacity available |
result = client.embeddings.create_and_wait(texts, content_type="topic", priority="critical")
Embedding Models
Supported Models
| Model | Provider | Dimensions | Custom Dims |
|---|---|---|---|
gemini-embedding-001 |
3072 | No | |
text-embedding-004 |
768 | No | |
text-multilingual-embedding-002 |
768 | No | |
text-embedding-3-small |
OpenAI | 1536 | Yes |
text-embedding-3-large |
OpenAI | 3072 | Yes |
Using a Specific Model
result = client.embeddings.create_and_wait(
texts=[{"id": "doc1", "text": "Hello world"}],
content_type="document",
embedding_model="text-embedding-3-small",
embedding_dimensions=512, # Custom dimensions (only for models that support it)
)
Content Hash
The SDK provides deterministic content hashing for learning tools.
from vector_sdk import compute_content_hash, extract_tool_text
# Compute hash for a FlashCard
hash = compute_content_hash(
"FlashCard",
{"type": "BASIC", "term": "Mitochondria", "definition": "The powerhouse of the cell"}
)
# Extract text for embedding
text = extract_tool_text(
"FlashCard",
{"type": "BASIC", "term": "Mitochondria", "definition": "The powerhouse of the cell"}
)
Migration from EmbeddingClient
The SDK now uses a namespace-based API with VectorClient. The old EmbeddingClient is preserved for backward compatibility.
Method Mapping
| Old (EmbeddingClient) | New (VectorClient) |
|---|---|
submit() |
client.embeddings.create() |
wait_for_result() |
client.embeddings.wait_for() |
submit_and_wait() |
client.embeddings.create_and_wait() |
get_queue_depth() |
client.embeddings.get_queue_depth() |
query() |
client.search.query() |
wait_for_query_result() |
client.search.wait_for() |
query_and_wait() |
client.search.query_and_wait() |
lookup_by_ids() |
client.db.get_by_ids() |
search_by_metadata() |
client.db.find_by_metadata() |
clone_from_namespace() |
client.db.clone() |
delete_from_namespace() |
client.db.delete() |
Migration Example
# Old API (still works, emits deprecation warnings)
from vector_sdk import EmbeddingClient
client = EmbeddingClient("redis://localhost:6379")
result = client.submit_and_wait(texts, content_type)
client.close()
# New API (recommended)
from vector_sdk import VectorClient
client = VectorClient(redis_url="redis://localhost:6379")
result = client.embeddings.create_and_wait(texts, content_type)
client.close()
Error Handling
from vector_sdk import VectorClient, ModelValidationError
try:
with VectorClient(redis_url="redis://localhost:6379") as client:
result = client.embeddings.create_and_wait(
texts=[{"id": "doc1", "text": "Hello"}],
content_type="test",
embedding_model="text-embedding-3-small",
timeout=30,
)
if result.is_success:
print("Success!")
elif result.is_partial:
print("Partial success. Errors:")
for err in result.errors:
print(f" - {err.id}: {err.error}")
except ModelValidationError as e:
print(f"Model validation failed: {e}")
except TimeoutError as e:
print(f"Request timed out: {e}")
except ValueError as e:
print(f"Invalid input: {e}")
Best Practices
1. Use Appropriate Priority
# Use appropriate priority levels
client.embeddings.create(texts, content_type="backfill", priority="low")
client.embeddings.create(texts, content_type="userRequest", priority="critical")
2. Batch Your Requests
# Batch multiple texts per request for efficiency
texts = [{"id": doc.id, "text": doc.text} for doc in documents]
client.embeddings.create(texts, content_type)
3. Use Context Managers
with VectorClient(redis_url="redis://...") as client:
# Client automatically closed on exit
pass
4. Deduplication
The gateway automatically deduplicates embedding requests using the contentHash metadata field. If a vector with the same contentHash already exists in the target namespace, the embedding generation is skipped to reduce costs.
- Structured embeddings: Deduplication is enabled by default for all tool types except Topics (which always re-embed since content may change for the same ID).
- Raw embeddings: Pass
allow_duplicates=Trueto skip deduplication when needed.
# Default: deduplication enabled (contentHash checked before embedding)
client.embeddings.create(texts, content_type="flashcard", storage=storage_config)
# Opt out of deduplication
client.embeddings.create(texts, content_type="topic", allow_duplicates=True)
The EmbeddingResult includes a skipped_count field showing how many items were deduplicated:
result = client.embeddings.create_and_wait(texts, content_type="flashcard")
print(f"Processed: {result.processed_count}, Skipped: {result.skipped_count}")
License
Proprietary - All rights reserved.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sf_vector_sdk-0.3.1b15.tar.gz.
File metadata
- Download URL: sf_vector_sdk-0.3.1b15.tar.gz
- Upload date:
- Size: 49.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.10.0 {"installer":{"name":"uv","version":"0.10.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2b7f18fe04cd600bd7a9ba152b65c8da2e056f10db53a135ad3b45a9e1f3ae89
|
|
| MD5 |
fd8c064424b18c4cb0f12f9eb4c15ad1
|
|
| BLAKE2b-256 |
42f400346227aa11ead48c9b192062e33723b61897651a37ad5ff6f64ad63aba
|
File details
Details for the file sf_vector_sdk-0.3.1b15-py3-none-any.whl.
File metadata
- Download URL: sf_vector_sdk-0.3.1b15-py3-none-any.whl
- Upload date:
- Size: 52.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.10.0 {"installer":{"name":"uv","version":"0.10.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1e8144183a1392508e6497f478fb05d82163a8c216e9cfb8022d39085e92934b
|
|
| MD5 |
27038405714f525b7efeb6aa89040576
|
|
| BLAKE2b-256 |
fd80ef885cb504c31c379440a0839d37e3477cee6adbe72beda94c986845d22d
|