Skip to main content

Scalable Objects Persistence (SOP) V2 for Python. General Public Availability (GPA) Release

Project description

SOP for Python (sop4py)

Scalable Objects Persistence (SOP) is a high-performance, transactional storage engine for Python, powered by a robust Go backend. It combines the raw speed of direct disk I/O with the reliability of ACID transactions and the flexibility of modern AI data management.

Key Features

  • Unified Database: Single entry point for managing Vector, Model, and Key-Value stores.
  • Transactional B-Tree Store: Unlimited, persistent B-Tree storage for key-value data.
  • Complex Keys: Support for composite keys (structs/dataclasses) with custom index specifications (e.g., Region -> Dept -> ID).
  • Metadata "Ride-on" Keys: Store metadata directly in the B-Tree key (e.g., timestamps, status flags) to enable high-speed scanning and filtering of millions of records without fetching the heavy value payload. Ideal for "Big Data" management and analytics.
  • Vector Database: Built-in vector search (k-NN) for AI embeddings and similarity search.
  • Text Search: Transactional, embedded text search engine (BM25).
  • AI Model Store: Versioned storage for machine learning models (B-Tree backed).
  • ACID Compliance: Full transaction support (Begin, Commit, Rollback) with isolation.
  • High Performance: Written in Go with a lightweight Python wrapper (ctypes).
  • Caching: Integrated Redis-backed L1/L2 caching for speed.
  • Replication & Fault Tolerance: Supports Erasure Coding for Blob Store (managing B-Tree nodes & large data files) to distribute data across drives with configurable parity. Also features Active/Passive Replication for the Registry to ensure high availability.
  • Multi-Tenancy: Native support for Cassandra Keyspaces or Directory-based isolation.
  • Flexible Deployment: Supports both Standalone (local) and Clustered (distributed) modes.

Performance & Big Data Management

SOP is designed for high-throughput, low-latency scenarios, making it suitable for "Big Data" management on commodity hardware.

  • "Ride-on" Metadata: By embedding metadata (like IsDeleted, LastUpdated, Category) directly into the Key struct but excluding it from the index (using IndexSpecification), you can scan millions of keys per second to filter data. This avoids the I/O penalty of fetching the full Value (which might be a large JSON blob or binary file) just to check a status flag.
  • Direct I/O: SOP bypasses OS page caches where appropriate to offer consistent, raw disk performance.
  • Parallelism: The underlying Go engine utilizes highly concurrent goroutines for managing B-Tree nodes and vector indexes.

Documentation

  • API Cookbook: Common recipes and patterns (Key-Value, Transactions, AI).
  • Examples: Complete runnable scripts.

Installation

Install directly from PyPI:

pip install sop4py

Data Browser & Full Data Management

SOP includes a powerful Data Browser that provides full data management capabilities for your B-Tree stores. It goes beyond simple viewing, offering a complete GUI for inspecting, searching, and manage your data at scale.

To launch it, simply run:

sop-httpserver

Key Capabilities

  • Full Data Management: Perform comprehensive CRUD (Create, Read, Update, Delete) operations on any record directly from the UI.
  • High-Performance Search: Utilizes B-Tree positioning for instant lookups, even in datasets with millions of records. Supports both simple keys and complex composite keys (e.g., searching by Country + City).
  • Efficient Navigation: Smart pagination and traversal controls (First, Previous, Next, Last) allow you to browse massive datasets without performance penalties.
  • Bulk Operations: Designed for rapid-fire management of records with a clean, non-distracting interface.
  • Responsive & Cross-Platform: Works seamlessly across diverse monitor sizes and devices.
  • Automatic Setup: The tool automatically downloads the correct binary for your OS/Architecture upon first run.

Usage: By default, it opens on http://localhost:8080. Arguments: You can pass standard flags, e.g., sop-httpserver -port 9090 -registry ./my_data.

Generating Sample Data

To see the Data Browser in action, you can generate a sample database with complex keys using the included example script:

  1. Run the generator:

    # If installed via pip
    sop-demo run large_complex_demo
    
    # Or manually if you have the source
    python3 examples/large_complex_demo.py
    

    This will create a database in data/large_complex_db with two stores: people (Complex Key) and products (Composite Key).

  2. Open in Browser:

    sop-httpserver -registry data/large_complex_db
    

Prerequisites

  • Redis: Required for caching and transaction coordination (especially in Clustered mode). Note: Redis is NOT used for data storage, just for coordination & to offer built-in caching.

Running the Examples

SOP comes with a bundled CLI tool sop-demo to easily list and run examples directly from your installation.

List available examples:

sop-demo list

Run a specific example:

sop-demo run vector_demo

Copy examples to your workspace: If you want to inspect the code or modify the examples, you can copy them to your local directory:

sop-demo copy
# Copies to ./sop_examples/

Manual Execution: If you have copied the examples locally, you can also run them using python directly:

python3 sop_examples/concurrent_demo.py

Concurrent Transactions (Standalone): This demo shows how to run concurrent transactions without a Redis dependency. It simulates real-world scenarios by introducing a small random sleep interval (jitter) between batch transactions to mimic network latency and reduce contention.

sop-demo run concurrent_demo_standalone

Concurrent Transactions (Clustered): This demo shows how to run concurrent transactions in a distributed environment (requires Redis). Similar to the standalone demo, it uses jitter to simulate realistic commit timing across different machines in a cluster.

sop-demo run concurrent_demo

Vector Search:

sop-demo run vector_demo

See the examples/ directory for more scripts. ```

  1. Set PYTHONPATH:
    export PYTHONPATH=$PYTHONPATH:$(pwd)/jsondb/python
    

Quick Start Guide

SOP uses a unified Database object to manage all types of stores (Vector, Model, and B-Tree). All operations are performed within a Transaction.

1. Initialize Database & Context

First, create a Context and open a Database connection.

from sop import Context, TransactionMode, TransactionOptions, Btree, BtreeOptions, Item
from sop.ai import Database, DatabaseType, Item as VectorItem
from sop.database import DatabaseOptions

# Initialize Context
ctx = Context()

# Open Database (Standalone Mode)
# This creates/opens a database at the specified path.
db = Database(DatabaseOptions(stores_folders=["data/my_db"], type=DatabaseType.Standalone))

# Open Database (Clustered Mode with Multi-Tenancy)
# Connects to a specific Cassandra Keyspace ("tenant_1").
# Requires Cassandra and Redis.
# db_clustered = Database(DatabaseOptions(stores_folders=["data/blobs"], keyspace="tenant_1", type=DatabaseType.Clustered))

2. Start a Transaction

All data operations (Create, Read, Update, Delete) must happen within a transaction.

# Begin a transaction (Read-Write)
# You can use 'with' block for auto-commit/rollback, or manage manually.
with db.begin_transaction(ctx) as tx:
    
    # --- 3. Vector Store (AI) ---
    # Open a Vector Store named "products"
    vector_store = db.open_vector_store(ctx, tx, "products")
    
    # Upsert a Vector Item
    vector_store.upsert(ctx, VectorItem(
        id="prod_101",
        vector=[0.1, 0.5, 0.9],
        payload={"name": "Laptop", "price": 999}
    ))

    # --- 4. Model Store (AI) ---
    # Open a Model Store named "classifiers"
    model_store = db.open_model_store(ctx, tx, "classifiers")
    
    # Save a Model
    model_store.save(ctx, "churn", "v1.0", {
        "algorithm": "random_forest",
        "trees": 100
    })

    # --- 5. B-Tree Store (Key-Value) ---
    # Open a B-Tree named "users"
    # Use new_btree to create a new store, or open_btree for existing ones.
    # BtreeOptions.name is optional if you pass the name directly to new_btree.
    btree = db.new_btree(ctx, "users", tx)
    
    # Add a Key-Value pair
    btree.add(ctx, Item(key="user_123", value="John Doe"))
    
    # Find a value
    if btree.find(ctx, "user_123"):
        # Fetch the value
        items = btree.get_values(ctx, Item(key="user_123"))
        if items and items[0].value:
            print(f"Found User: {items[0].value}")

    # --- 6. Complex Keys (Structs) ---
    # Define a composite key using a dataclass
    from dataclasses import dataclass
    from sop.btree import IndexSpecification, IndexFieldSpecification

    @dataclass
    class EmployeeKey:
        region: str
        department: str
        id: int

    # Create B-Tree with custom index (Region -> Dept -> ID)
    # This enables fast prefix scans (e.g., "Get all employees in US")
    spec = IndexSpecification(index_fields=(
        IndexFieldSpecification("region", ascending_sort_order=True),
        IndexFieldSpecification("department", ascending_sort_order=True),
        IndexFieldSpecification("id", ascending_sort_order=True)
    ))
    
    # Pass spec as index_spec argument
    employees = db.new_btree(ctx, "employees", tx, index_spec=spec)

    # Add item with complex key
    employees.add(ctx, Item(
        key=EmployeeKey("US", "Sales", 101), 
        value={"name": "Alice"}
    ))

    # --- 7. Simplified Lookup (Dictionary Keys) ---
    # You can search for items using a plain dictionary, without needing the original dataclass.
    # This is useful for consumer apps that just need to read data.
    
    # Open existing B-Tree (no IndexSpec needed, it's loaded from disk)
    employees_read = db.open_btree(ctx, "employees", tx)
    
    # Search using a dict matching the key structure
    if employees_read.find(ctx, {"region": "US", "department": "Sales", "id": 101}):
        print("Found Alice!")

    # --- 8. Text Search ---
    # Open a Search Index
    idx = db.open_search(ctx, "articles", tx)
    idx.add("doc1", "The quick brown fox")

# Transaction commits automatically here.
# If an exception occurs, it rolls back.

6. Querying Data

You can perform queries in a separate transaction (e.g., Read-Only).

# Begin a Read-Only transaction (optional optimization)
with db.begin_transaction(ctx, mode=TransactionMode.ForReading.value) as tx:
    
    # --- Vector Search ---
    vs = db.open_vector_store(ctx, tx, "products")
    hits = vs.query(ctx, vector=[0.1, 0.5, 0.8], k=5)
    for hit in hits:
        print(f"Vector Match: {hit.id}, Score: {hit.score}")

    # --- Model Retrieval ---
    ms = db.open_model_store(ctx, tx, "classifiers")
    model = ms.get(ctx, "churn", "v1.0")
    print(f"Loaded Model: {model['algorithm']}")

    # --- B-Tree Lookup ---
    us = db.open_btree(ctx, "user_store", tx)
    if us.find(ctx, "user1"):
        # Fetch the current item
        item = us.get_current_item(ctx)
        print(f"User Found: {item.value}")

Performance Tip: For Vector Search workloads that are "Build-Once-Query-Many", use TransactionMode.NoCheck. This bypasses transaction overhead for maximum query throughput.

# High-performance Vector Search (No ACID checks)
with db.begin_transaction(ctx, mode=TransactionMode.NoCheck.value) as tx:
    vs = db.open_vector_store(ctx, tx, "products")
    hits = vs.query(ctx, vector=[0.1, 0.5, 0.8], k=5)

Advanced Configuration

Logging

You can configure the internal logging of the SOP engine (Go backend) to output to a file or standard error, and control the verbosity.

from sop import Logger, LogLevel

# Configure logging to a file with Debug level
Logger.configure(LogLevel.Debug, "sop_engine.log")

# Or configure logging to stderr (default) with Info level
Logger.configure(LogLevel.Info)

Transaction Options

You can configure timeouts, isolation levels, and more.

from sop import TransactionOptions

opts = TransactionOptions(
    max_time=15,  # 15 minutes timeout
)

tx = db.begin_transaction(ctx, options=opts)

Clustered Mode

For distributed deployments, switch to DatabaseType.Clustered. This requires Redis for coordination.

from sop.ai import DatabaseType

db = Database(
    ctx, 
    stores_folders=["/mnt/shared_data"], 
    type=DatabaseType.Clustered
)

Clustered Backend Setup (Cassandra + Redis)

For production environments using Clustered mode, you should initialize both Cassandra (for storage) and Redis (for distributed locking and caching) at application startup.

from sop import Redis
from sop.cassandra import Cassandra
from sop.database import Database, DatabaseOptions, DatabaseType

# 1. Initialize Redis (Required for Locking/Caching in Clustered mode)
# Format: redis://<user>:<password>@<host>:<port>/<db_number>
Redis.initialize("redis://:password@localhost:6379/0")

# 2. Initialize Cassandra (Global Connection)
Cassandra.initialize({
    "cluster_hosts": ["127.0.0.1"],
    "consistency": 1,          # 1 = LocalQuorum
    "authenticator": {
        "username": "cassandra",
        "password": "password"
    }
})

# ... Application Logic ...

# Connect to a specific tenant's keyspace
db = Database(DatabaseOptions(
    keyspace="tenant_1",
    type=DatabaseType.Clustered
))

# ...

# Cleanup on shutdown
Redis.close()
Cassandra.close()

Architecture

SOP uses a split architecture:

  1. Core Engine (Go): Handles disk I/O, B-Tree algorithms, caching, and transactions. Compiled as a shared library (.dylib, .so, .dll).
  2. Python Wrapper: Uses ctypes to interface with the Go engine, providing a Pythonic API (sop package).

Project Links

Contributing

Contributions are welcome! Please check the CONTRIBUTING.md file in the repository for guidelines.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sop4py-2.0.42.tar.gz (29.3 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

sop4py-2.0.42-py3-none-win_amd64.whl (7.6 MB view details)

Uploaded Python 3Windows x86-64

sop4py-2.0.42-py3-none-manylinux_2_17_x86_64.whl (7.6 MB view details)

Uploaded Python 3manylinux: glibc 2.17+ x86-64

sop4py-2.0.42-py3-none-manylinux_2_17_aarch64.whl (6.9 MB view details)

Uploaded Python 3manylinux: glibc 2.17+ ARM64

sop4py-2.0.42-py3-none-macosx_11_0_arm64.whl (3.8 MB view details)

Uploaded Python 3macOS 11.0+ ARM64

sop4py-2.0.42-py3-none-macosx_10_9_x86_64.whl (4.2 MB view details)

Uploaded Python 3macOS 10.9+ x86-64

File details

Details for the file sop4py-2.0.42.tar.gz.

File metadata

  • Download URL: sop4py-2.0.42.tar.gz
  • Upload date:
  • Size: 29.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for sop4py-2.0.42.tar.gz
Algorithm Hash digest
SHA256 17e061ab02eaf87f9846c9f291023c53a1e050b315d598eb1cd2cff6fae0cbd4
MD5 2c7e4262949eb89fd956427c8d35ebbc
BLAKE2b-256 30185e50f65165f789b6f7317fb70ef76469339eddfbb5e41a5f05eabe383172

See more details on using hashes here.

File details

Details for the file sop4py-2.0.42-py3-none-win_amd64.whl.

File metadata

  • Download URL: sop4py-2.0.42-py3-none-win_amd64.whl
  • Upload date:
  • Size: 7.6 MB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for sop4py-2.0.42-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 ca8372e0e5a8b967e234e4c2ef4cb2e70ddedbf4784670a42ebb9ce519c29be9
MD5 748ab67fc0e005cb4213cd8a87200159
BLAKE2b-256 b6a3171dcf6f8c020d8566543c622279d160541da9e19680d75ec02f5ed35f0e

See more details on using hashes here.

File details

Details for the file sop4py-2.0.42-py3-none-manylinux_2_17_x86_64.whl.

File metadata

File hashes

Hashes for sop4py-2.0.42-py3-none-manylinux_2_17_x86_64.whl
Algorithm Hash digest
SHA256 7b57334a69b5c94e8136246240295b33eb9ad0d2a342c0042e78fd49146f17c7
MD5 480a69e8a8e0dc63d10d6fafc2ab313b
BLAKE2b-256 6ca8ce7371a80d6f85ea862d4841db8a8c8612d45e1dfe8da71116e69ce4ae4e

See more details on using hashes here.

File details

Details for the file sop4py-2.0.42-py3-none-manylinux_2_17_aarch64.whl.

File metadata

File hashes

Hashes for sop4py-2.0.42-py3-none-manylinux_2_17_aarch64.whl
Algorithm Hash digest
SHA256 95dbc9e0b48aedb3179eb0340cc76ce90511d9a1b4a5a3566599a05f35d91ff4
MD5 ee6f4fdacdf418e04b89ebd616223f06
BLAKE2b-256 2236130867e56b43ce13d4072016a5256562d8e104245c69b3e5ed586c85e977

See more details on using hashes here.

File details

Details for the file sop4py-2.0.42-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for sop4py-2.0.42-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 46e52b334786a568b16e666a0b15238e91f8cd65690872b49789d557e2917ff0
MD5 23c9aa8b8aa7a5515a1425ff405089d2
BLAKE2b-256 a8ccb7c702b6846ec0100c7e70c1b4ba09204b0b0f4c7d7a4e4ba8e39561f3c1

See more details on using hashes here.

File details

Details for the file sop4py-2.0.42-py3-none-macosx_10_9_x86_64.whl.

File metadata

File hashes

Hashes for sop4py-2.0.42-py3-none-macosx_10_9_x86_64.whl
Algorithm Hash digest
SHA256 d9d13cbdb26f10c701f3be686bc93b9eb403e2b17989449fbb882a35957097ca
MD5 c80c023c25160c40d297aa340c295309
BLAKE2b-256 4d372724ecbfe7a540a3c6d2345bd53b79099b3d7e863f3563942bc4097693ba

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page