Skip to main content

cqlite-py

Python bindings for CQLite - a high-performance library for reading Apache Cassandra 5.0 SSTable files locally, without requiring a running Cassandra cluster.

Installation

pip install cqlite-py

Quick Start

import cqlite

# Open a database with schema
with cqlite.open('path/to/sstables', schema='schema.cql') as db:
    # Execute queries
    for row in db.execute('SELECT * FROM keyspace.table LIMIT 10'):
        print(row.to_dict())

Features

  • Zero cluster dependency - Read SSTable files directly from disk
  • Full CQL type support - All primitive types, collections, UDTs, and frozen types
  • Memory-efficient streaming - Iterate over large datasets without loading all rows
  • Thread-safe database handles - Safe concurrent access from multiple threads
  • Cross-platform - Linux (x86_64, ARM64), macOS (Intel, Apple Silicon), Windows

Supported Platforms

Platform Architecture Status
Linux x86_64
Linux ARM64
macOS Intel (x86_64)
macOS Apple Silicon
Windows x64

Requirements

  • Python 3.9+
  • Cassandra 5.0 SSTable files

API Reference

Opening a Database

import cqlite

# Context manager (recommended)
with cqlite.open(data_dir, schema=schema_path) as db:
    # use db...

# Manual management
db = cqlite.open(data_dir, schema=schema_path)
# use db...
db.close()

Executing Queries

# Simple query
results = db.execute('SELECT * FROM keyspace.table')
for row in results:
    print(row.to_dict())

# With LIMIT
for row in db.execute('SELECT name, age FROM users LIMIT 100'):
    print(f"{row['name']}: {row['age']}")

# Access query metadata
print(f"Rows returned: {len(results)}")
print(f"Execution time: {results.execution_time_ms}ms")
print(f"Columns: {[col.name for col in results.columns]}")

Streaming Large Results

For memory-efficient iteration over large datasets:

from cqlite import StreamingConfig

# Configure streaming for memory efficiency
config = StreamingConfig(buffer_size=512, chunk_size=1000)
for row in db.execute_streaming('SELECT * FROM large_table', config=config):
    process(row)

# Track progress
iterator = db.execute_streaming('SELECT * FROM large_table')
for row in iterator:
    if iterator.rows_received % 10000 == 0:
        print(f"Processed {iterator.rows_received} rows")

Configuration Presets

import cqlite

# Built-in presets for common use cases
config = cqlite.memory_optimized()      # 256 MB max memory
config = cqlite.performance_optimized() # 4 GB max memory

# Open database with preset configuration
db = cqlite.open('path/to/data', schema='schema.cql', config='memory_optimized')

# Validate custom configuration
custom_config = {'memory': {'max_memory': 536870912}}  # 512 MB
cqlite.validate_config(custom_config)

Refreshing SSTables (v0.13)

If Cassandra (or another process) writes new SSTables while your database handle is open, call refresh() to re-discover them. Refresh is explicit-only (CQLite never rescans behind your back) and atomic / fail-closed: if any newly found generation fails to open, the swap is rolled back and the handle keeps serving the prior, consistent set of readers.

import cqlite

with cqlite.open('path/to/sstables', schema='schema.cql') as db:
    # ... time passes; Cassandra flushes/compacts new SSTables to disk ...

    report = db.refresh()
    print(f'Tables scanned:  {report.tables_scanned}')
    print(f'Readers added:   {report.readers_added}')
    print(f'Readers removed: {report.readers_removed}')

    # Subsequent queries see the newly discovered data
    for row in db.execute('SELECT * FROM keyspace.table'):
        print(row.to_dict())

refresh() returns a RefreshReport with the integer attributes tables_scanned, readers_added, and readers_removed, plus a to_dict() helper. It raises RuntimeError if the database is already closed, and CqliteError if a newly discovered generation fails to open (the prior reader set is preserved).

Result Byte Budget (v0.13)

Non-streaming execute() queries are bounded by a result-size budget, defaulting to 64 MiB (64 * 1024 * 1024 bytes). When the materialized result's running byte estimate exceeds the budget, the query fails with a cqlite.QueryError directing you to add a LIMIT clause or switch to execute_streaming(). Streaming queries are not subject to this budget.

Adjust the budget with the max_result_bytes key in the config passed to cqlite.open() (absent, it stays at 64 MiB):

import cqlite

# Raise the budget to 256 MiB for this handle
config = {'max_result_bytes': 256 * 1024 * 1024}

try:
    with cqlite.open('path/to/data', schema='schema.cql', config=config) as db:
        rows = db.execute('SELECT * FROM keyspace.big_table')
        for row in rows:
            process(row)
except cqlite.QueryError as e:
    # Result exceeded the byte budget — add a LIMIT or stream instead
    print(f'Result too large: {e}')
    for row in db.execute_streaming('SELECT * FROM keyspace.big_table'):
        process(row)

OpenTelemetry Tracing (v0.13)

CQLite can emit OpenTelemetry traces when built with the observability Cargo feature; without that feature the configuration is accepted but is a no-op. Pass an otel_config dict to cqlite.open(). Values are layered over the CQLITE_OTEL_* environment variables, and OpenTelemetry is initialized once per process.

import cqlite

db = cqlite.open(
    'path/to/sstables',
    schema='schema.cql',
    otel_config={
        'enabled': True,                       # default False
        'endpoint': 'http://localhost:4317',   # default 'http://localhost:4317'
        'protocol': 'grpc',                    # 'grpc' (default) or 'http'
        'service_name': 'cqlite',              # default 'cqlite'
        'service_version': '0.13.0',           # default: package version
        'sampling_ratio': 1.0,                 # default 1.0
        'timeout_ms': 10000,                   # default 10000
    },
)

Unknown keys raise ValueError.

Error Handling

import cqlite

try:
    with cqlite.open('path/to/data', schema='schema.cql') as db:
        result = db.execute('SELECT * FROM keyspace.table')
        for row in result:
            print(row.to_dict())
except cqlite.ParseError as e:
    print(f"Query syntax error: {e}")
except cqlite.QueryError as e:
    print(f"Query execution failed: {e}")
except cqlite.SchemaError as e:
    print(f"Schema validation failed: {e}")
except IOError as e:
    print(f"File not found: {e}")
except RuntimeError as e:
    print(f"Database already closed: {e}")

Exception Hierarchy:

CqliteError (base exception)
├── SchemaError   - Schema parsing or validation failures
├── QueryError    - Query execution failures
└── ParseError    - CQL syntax errors

Built-in exceptions also used:
├── IOError       - File system errors
├── ValueError    - Invalid configuration
├── RuntimeError  - Invalid state (e.g., database closed)
└── MemoryError   - Memory allocation failures

Type Conversions

CQL types are automatically converted to Python native types:

CQL Type Python Type
text, varchar str
int, bigint, smallint, tinyint int
float, double float
boolean bool
blob bytes
timestamp datetime.datetime
date datetime.date
time int (nanoseconds since midnight, lossless)
duration cqlite.Duration (exact months / days / nanos)
uuid, timeuuid uuid.UUID
inet ipaddress.IPv4Address or IPv6Address
decimal decimal.Decimal
varint int (arbitrary precision)
list<T> list
set<T> frozenset
map<K,V> dict
tuple<...> tuple
frozen<T> Unwrapped inner type
UDT dict with _type and _keyspace keys

Write Operations

CQLite v0.9.0 adds write support to the Python bindings. Open the database with writable=True and a write_dir to enable write operations.

import cqlite

with cqlite.open(
    'path/to/sstables',
    schema='schema.cql',
    writable=True,
    write_dir='/tmp/my-writes',
) as db:
    # Write rows via CQL INSERT, UPDATE, or DELETE
    db.execute(
        "INSERT INTO test_basic.simple_table (id, name, age) "
        "VALUES (11111111-1111-1111-1111-111111111111, 'Alice', 30)"
    )
    db.execute(
        "UPDATE test_basic.simple_table SET age = 31 "
        "WHERE id = 11111111-1111-1111-1111-111111111111"
    )

    # Flush the in-memory write buffer (memtable) to an SSTable on disk.
    # Returns the path to the flushed Data.db file.
    path = db.flush_run()
    print(f'Flushed to: {path}')

    # Run background compaction within a time budget
    report = db.maintenance_step(budget_ms=100)
    print(f'Merged {report.rows_merged} rows in {report.time_spent_ms:.1f} ms')
    if report.pending_compaction:
        print('More compaction work available')

    # Inspect write statistics
    stats = db.write_stats
    print(f'Memtable size: {stats.memtable_size_bytes} bytes')
    print(f'Total flushed: {stats.total_written_bytes} bytes')

Write API

Method / Property Description
db.execute(cql) Execute a CQL INSERT, UPDATE, or DELETE statement
db.flush_run() Flush memtable to SSTable; returns the Data.db path or "" if memtable was empty
db.maintenance_step(budget_ms) Run STCS compaction for up to budget_ms milliseconds; returns MaintenanceReport
db.write_stats WriteStats property: memtable_size_bytes, memtable_row_count, total_written_bytes, l0_sstable_count

Known Limitations

  • Counter columns cannot be written — execute() raises CqliteError for counter mutations.
  • Concurrent queries on the same handle may need a warm-up query first (Issue #311).

See docs/write-support-limitations.md for the full limitations reference.

Resources

License

MIT OR Apache-2.0

Links

Release files for cqlite-py 0.16.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cqlite-py 0.16.0
File Size Uploaded
cqlite_py-0.16.0.tar.gz 4.8 MB Details

Built distributions (wheels)

Table of built distributions (wheels) for cqlite-py 0.16.0
File
cqlite_py-0.16.0-cp39-abi3-win_amd64.whl CPython 3.9 abi3 Windows x86-64 Details
cqlite_py-0.16.0-cp39-abi3-manylinux_2_28_x86_64.whl CPython 3.9 abi3 Linux glibc 2.28+ x86-64 Details
cqlite_py-0.16.0-cp39-abi3-manylinux_2_28_aarch64.whl CPython 3.9 abi3 Linux glibc 2.28+ ARM64 Details
cqlite_py-0.16.0-cp39-abi3-macosx_11_0_arm64.whl CPython 3.9 abi3 macOS 11.0+ ARM64 Details
cqlite_py-0.16.0-cp39-abi3-macosx_10_12_x86_64.whl CPython 3.9 abi3 macOS 10.12+ x86-64 Details

Total release size: 32.7 MB

Release files / cqlite_py-0.16.0.tar.gz

Download URL cqlite_py-0.16.0.tar.gz
Size 4.8 MB
Tags Source
SHA-256 checksum
How to use checksums
047f2b0999cef96bb729ae900ad4c8547bd6a2b3cba3964441ac21505c960b92
BLAKE2b-256 checksum
How to use checksums
ba522f83d2e0dfc8e73b20a0e87bd24a54a732b77ef59696e3adf2d9a3a671cd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release files / cqlite_py-0.16.0-cp39-abi3-win_amd64.whl

Download URL cqlite_py-0.16.0-cp39-abi3-win_amd64.whl
Size 6.0 MB
Tags CPython 3.9 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
72692c6528a17748f673f082d430fb1b0241e6531a3ac6ff0518ae7129682456
BLAKE2b-256 checksum
How to use checksums
e2672a62bc48fda8adaa5c67f05cae6154868dd1e54134735e04b2aef3e368c7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release files / cqlite_py-0.16.0-cp39-abi3-manylinux_2_28_x86_64.whl

Download URL cqlite_py-0.16.0-cp39-abi3-manylinux_2_28_x86_64.whl
Size 5.8 MB
Tags CPython 3.9 Linux glibc 2.28+ x86-64 abi3
SHA-256 checksum
How to use checksums
bf108e3e9db5c14466dd84dd344dcbf36eae9024d2c3241bb1261eb89a6d26c1
BLAKE2b-256 checksum
How to use checksums
e63403cb8150c9f499972059c4f4d803dc3019ef7f4e46765dff8279c6adbdd2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release files / cqlite_py-0.16.0-cp39-abi3-manylinux_2_28_aarch64.whl

Download URL cqlite_py-0.16.0-cp39-abi3-manylinux_2_28_aarch64.whl
Size 5.4 MB
Tags CPython 3.9 Linux glibc 2.28+ ARM64 abi3
SHA-256 checksum
How to use checksums
7211b6b7782544de061d86e5d82520f59b47479341256bb578447106ece3f418
BLAKE2b-256 checksum
How to use checksums
5639a9793fdc0039955948193b3702b35c05a3235da2f79759f8f71b6633ecaf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release files / cqlite_py-0.16.0-cp39-abi3-macosx_11_0_arm64.whl

Download URL cqlite_py-0.16.0-cp39-abi3-macosx_11_0_arm64.whl
Size 5.1 MB
Tags CPython 3.9 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
d6c394907ecff8790f237319c53f3d3311a1623eb1b27a9bbe56105c9875d7c2
BLAKE2b-256 checksum
How to use checksums
c14d6aebab87281b9c9ecc49c36ebbd894484d8a175e0815109e6d647890b174
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release files / cqlite_py-0.16.0-cp39-abi3-macosx_10_12_x86_64.whl

Download URL cqlite_py-0.16.0-cp39-abi3-macosx_10_12_x86_64.whl
Size 5.6 MB
Tags CPython 3.9 abi3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
321849440918de4a2446001a87c804ece986a788c6db9b2d81391849cda9a86a
BLAKE2b-256 checksum
How to use checksums
c75ff3f77a0fdc139a67a35c1ebc00c352a18ebe8f6cfd807c0101eb85f81a79
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release history Release notifications | RSS feed

0.16.1

6 release files

This release

0.16.0 This release

6 release files

0.15.0

6 release files

0.14.1

6 release files

0.14.0

6 release files

0.12.0

6 release files

0.11.0

6 release files

0.9.2

6 release files

0.9.1

6 release files

0.9.0

6 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page