Skip to main content

slick-queue-py

Python implementation of SlickQueue - a lock-free multi-producer multi-consumer (MPMC) queue with C++ interoperability through shared memory.

This is the Python binding for the SlickQueue C++ library. The Python implementation maintains exact binary compatibility with the C++ version, enabling seamless interprocess communication between Python and C++ applications.

License: MIT CI GitHub release

Features

  • Dual Mode Operation:
    • Local Memory Mode: In-process queue using local memory (no shared memory overhead)
    • Shared Memory Mode: Inter-process queue for interprocess communication
  • Lock-Free Multi-Producer Multi-Consumer: True MPMC support using atomic operations
  • C++/Python Interoperability: Python and C++ processes can share the same queue
  • Cross-Platform: Windows and Linux/macOS support (x86-64)
  • Memory Layout Compatible: Exact binary compatibility with C++ slick::queue<T>
  • Configurable Features: Optional behaviour is selected per queue through traits, so a feature you do not need costs nothing
  • High Performance: Hardware atomic operations for minimal overhead

Requirements

  • Python 3.8+ (uses multiprocessing.shared_memory)
  • 64-bit platform
  • For true lock-free operation: x86-64 CPU with CMPXCHG16B support (most CPUs since 2006)

Installation

pip install -e .

Or just copy the Python files to your project.

Quick Start

Local Memory Mode (Single Process)

from slick_queue_py import SlickQueue

# Create a queue in local memory (no shared memory)
q = SlickQueue(size=1024, element_size=256)

# Producer: Reserve a slot, write data, and publish
idx = q.reserve()
buf = q[idx]
buf[:len(b'hello')] = b'hello'
q.publish(idx)

# Consumer: Read data
read_index = 0
data, size, read_index = q.read(read_index)
if data is not None:
    print(f"Received: {data[:size]}")

q.close()  # unlink() does nothing for local mode

Shared Memory Mode (Multi-Process)

from slick_queue_py import SlickQueue

# Create a new shared memory queue (size must be power of two)
q = SlickQueue(name='my_queue', size=1024, element_size=256)

# Producer: Reserve a slot, write data, and publish
idx = q.reserve()
buf = q[idx]
buf[:len(b'hello')] = b'hello'
q.publish(idx)

# Consumer: Read data
read_index = 0
data, size, read_index = q.read(read_index)
if data is not None:
    print(f"Received: {data[:size]}")

q.close()
q.unlink()  # Delete shared memory segment

Multi-Producer Usage

from multiprocessing import Process
from slick_queue_py import SlickQueue
import struct

def producer_worker(queue_name, worker_id, num_items):
    # Open existing queue
    q = SlickQueue(name=queue_name, element_size=32)

    for i in range(num_items):
        # Reserve slot (thread-safe with atomic CAS)
        idx = q.reserve(1)

        # Write unique data
        data = struct.pack("<I I", worker_id, i)
        slot = q[idx]
        slot[:len(data)] = data

        # Publish (makes data visible to consumers)
        q.publish(idx, 1)

    q.close()

# Create queue
q = SlickQueue(name='mpmc_queue', size=64, element_size=32)

# Start multiple producers
producers = []
for i in range(4):
    p = Process(target=producer_worker, args=('mpmc_queue', i, 100))
    p.start()
    producers.append(p)

# Wait for completion
for p in producers:
    p.join()

q.close()
q.unlink()

Multi-Consumer Work-Stealing

For multiple consumers sharing work from a single queue, use an AtomicCursor to enable work-stealing patterns where each item is consumed by exactly one consumer.

Local Mode (Multi-Threading)

from threading import Thread
from slick_queue_py import SlickQueue, AtomicCursor
import struct

def consumer_worker(q, cursor, worker_id, results):
    items_processed = 0
    while True:
        # Atomically claim next item (work-stealing)
        data, size, index = q.read(cursor)

        if data is None:
            break  # No more data

        # Process the claimed item
        worker, seq = struct.unpack("<I I", data[:8])
        items_processed += 1

    results[worker_id] = items_processed

# Create local queue and cursor
q = SlickQueue(size=64, element_size=32)
cursor_buf = bytearray(8)
cursor = AtomicCursor(cursor_buf, 0)
cursor.store(0)  # Initialize cursor to 0

# Producer writes items
for i in range(100):
    idx = q.reserve()
    data = struct.pack("<I I", 0, i)
    q[idx][:len(data)] = data
    q.publish(idx)

# Start multiple consumer threads that share the work
results = {}
threads = []
for i in range(4):
    t = Thread(target=consumer_worker, args=(q, cursor, i, results))
    t.start()
    threads.append(t)

# Wait for all consumers
for t in threads:
    t.join()

print(f"Total items processed: {sum(results.values())}")
q.close()

Shared Memory Mode (Multi-Process)

from multiprocessing import Process, shared_memory
from slick_queue_py import SlickQueue, AtomicCursor
import struct

def consumer_worker(queue_name, cursor_name, worker_id):
    # Open shared queue and cursor
    q = SlickQueue(name=queue_name, element_size=32)
    cursor_shm = shared_memory.SharedMemory(name=cursor_name)
    cursor = AtomicCursor(cursor_shm.buf, 0)

    items_processed = 0
    while True:
        # Atomically claim next item (work-stealing)
        data, size, index = q.read(cursor)

        if data is None:
            break  # No more data

        # Process the claimed item
        worker, seq = struct.unpack("<I I", data[:8])
        items_processed += 1

    print(f"Worker {worker_id} processed {items_processed} items")
    cursor_shm.close()
    q.close()

# Create queue and shared cursor
q = SlickQueue(name='work_queue', size=64, element_size=32)
cursor_shm = shared_memory.SharedMemory(name='work_cursor', create=True, size=8)
cursor = AtomicCursor(cursor_shm.buf, 0)
cursor.store(0)  # Initialize cursor to 0

# Producer writes items
for i in range(100):
    idx = q.reserve()
    data = struct.pack("<I I", 0, i)
    q[idx][:len(data)] = data
    q.publish(idx)

# Start multiple consumer processes that share the work
consumers = []
for i in range(4):
    p = Process(target=consumer_worker, args=('work_queue', 'work_cursor', i))
    p.start()
    consumers.append(p)

# Wait for all consumers
for p in consumers:
    p.join()

cursor_shm.close()
cursor_shm.unlink()
q.close()
q.unlink()

C++/Python Interoperability

The Python implementation is fully compatible with the C++ SlickQueue library. Python and C++ processes can produce and consume from the same queue with:

  • Exact memory layout compatibility: Binary-compatible with slick::queue<T>
  • Atomic operation compatibility: Same 16-byte and 8-byte CAS semantics
  • Bidirectional communication: C++ ↔ Python in both directions
  • Multi-producer support: Mix C++ and Python producers on the same queue

Platform Support for C++/Python Interop:

  • ✅ Linux/macOS: Full interoperability (both use POSIX shm_open)
  • ✅ Windows: Full interoperability
  • ✅ Python-only: Works on all platforms (Windows/Linux/macOS)

Basic C++ → Python Example

C++ Producer:

#include <slick/queue.hpp>

int main() {
    // Open existing queue created by Python
    slick::queue<uint8_t> q(32, "shared_queue");

    for (int i = 0; i < 100; i++) {
        auto idx = q.reserve();
        uint32_t value = i;
        std::memcpy(q[idx], &value, sizeof(value));
        q.publish(idx);
    }
}

Python Consumer:

from slick_queue_py import SlickQueue
import struct

# Create queue that C++ will write to
q = SlickQueue(name='shared_queue', size=64, element_size=32)

read_index = 0
for _ in range(100):
    data, size, read_index = q.read(read_index)
    if data is not None:
        value = struct.unpack("<I", data[:4])[0]
        print(f"Received from C++: {value}")

q.close()
q.unlink()

Building C++ Programs

To use the C++ SlickQueue library with your Python queues:

# Clone the C++ library
git clone https://github.com/SlickQuant/slick-queue.git

# Build your C++ program
g++ -std=c++17 -I slick-queue/include my_program.cpp -o my_program

Or use CMake (see CMakeLists.txt for reference):

include(FetchContent)
FetchContent_Declare(
    slick-queue
    GIT_REPOSITORY https://github.com/SlickQuant/slick-queue.git
    GIT_TAG main
)
FetchContent_MakeAvailable(slick-queue)

add_executable(my_program my_program.cpp)
target_link_libraries(my_program PRIVATE slick::queue)

See tests/test_interop.py and tests/cpp_*.cpp for comprehensive examples.

API Reference

SlickQueue

__init__(*, name=None, size=None, element_size=None, traits=None)

Create a queue in local memory or shared memory mode.

Parameters:

  • name (str, optional): Shared memory segment name. If None, uses local memory mode (single process).
  • size (int): Queue capacity (must be power of 2). Required for local mode or when creating shared memory.
  • element_size (int, required): Size of each element in bytes
  • traits (type, optional): Feature configuration, a QueueTraits subclass. Defaults to default_queue_traits. See Configuring Features (Traits).

Raises:

  • ValueError: If size is not a power of two
  • TypeError: If traits is missing a trait or declares one as something other than a bool
  • RuntimeError: If an existing segment disagrees about the layout marker - it carries none (created before slick-queue v1.4.0), was created with unknown layout features, or disagrees about enable_read_last

Examples:

# Local memory mode (single process)
q = SlickQueue(size=256, element_size=64)

# Create new shared memory queue
q = SlickQueue(name='my_queue', size=256, element_size=64)

# Open existing shared memory queue
q2 = SlickQueue(name='my_queue', element_size=64)

# Opt out of read_last() tracking to drop its CAS from publish()
class Lean(QueueTraits):
    enable_read_last = False

q3 = SlickQueue(size=256, element_size=64, traits=Lean)

reserve(n=1) -> int

Reserve n elements for writing. Multi-producer safe using atomic CAS.

Parameters:

  • n (int): Number of elements to reserve (default 1)

Returns:

  • int: Starting index of reserved space

Example:

idx = q.reserve(1)  # Reserve 1 elements

publish(index, n=1)

Publish data written to reserved space. Uses atomic operations with release memory ordering.

Parameters:

  • index (int): Index returned by reserve()
  • n (int): Number of elements to publish (default 1)

Example:

idx = q.reserve()
q[idx][:data_len] = data
q.publish(idx)

read(read_index) -> Tuple[Optional[bytes], int, int] or read(atomic_cursor) -> Tuple[Optional[bytes], int]

Read from queue with two modes:

Single-Consumer Mode (when read_index is int): Uses a plain int cursor for single-consumer scenarios. Returns the new read_index.

Multi-Consumer Mode (when read_index is AtomicCursor): Uses an atomic cursor for work-stealing/load-balancing across multiple consumers. Each consumer atomically claims items, ensuring each item is consumed exactly once.

Parameters:

  • read_index (int or AtomicCursor): Current read position or shared atomic cursor

Returns:

  • Single-consumer: Tuple[Optional[bytes], int, int] - (data or None, size, new_read_index)
  • Multi-consumer: Tuple[Optional[bytes], int] - (data or None, size)

API Difference from C++: Unlike C++ where read_index is updated by reference, the Python single-consumer version returns the new index. This is the Pythonic pattern since Python doesn't have true pass-by-reference.

# Python single-consumer (returns new index)
data, size, read_index = q.read(read_index)

# Python multi-consumer (atomic cursor)
from slick_queue_py import AtomicCursor
cursor = AtomicCursor(cursor_shm.buf, 0)
data, size, index = q.read(cursor)  # Atomically claim next item

# C++ (updates by reference for both)
auto [data, size] = queue.read(read_index);  // read_index modified in-place
auto [data, size] = queue.read(atomic_cursor);  // atomic_cursor modified in-place

Single-Consumer Example:

read_index = 0
while True:
    data, size, read_index = q.read(read_index)
    if data is not None:
        process(data)

Multi-Consumer Example (Local Mode - Threading):

from slick_queue_py import AtomicCursor

# Create local cursor for multi-threading
cursor_buf = bytearray(8)
cursor = AtomicCursor(cursor_buf, 0)
cursor.store(0)

# Multiple threads can share this cursor
while True:
    data, size, index = q.read(cursor)  # Each thread atomically claims items
    if data is not None:
        process(data)

Multi-Consumer Example (Shared Memory Mode - Multiprocess):

from multiprocessing import shared_memory
from slick_queue_py import AtomicCursor

# Create shared cursor for multi-process
cursor_shm = shared_memory.SharedMemory(name='cursor', create=True, size=8)
cursor = AtomicCursor(cursor_shm.buf, 0)
cursor.store(0)

# Multiple processes can share this cursor
while True:
    data, size, index = q.read(cursor)  # Each process atomically claims items
    if data is not None:
        process(data)

read_last() -> Tuple[Optional[bytes], int]

Read the most recently published item. Requires traits.enable_read_last.

Returns:

  • Tuple[Optional[bytes], int]: Tuple of (data, size)
    • data: Last published data, or None if the queue is empty or the slot was recycled by a wrapping producer while it was being read
    • size: Number of slots the item occupies (0 if no data is returned)

Raises:

  • RuntimeError: If traits.enable_read_last is False. There is no fallback: without the feature nothing maintains the last published index, and the reserved cursor is not a substitute because it reports reservations that were never published and truncates sizes above 65,535.

Example:

data, size = q.read_last()
if data is not None:
    print(f"Last item: {data[:size * element_size]}")

loss_count() -> int

Number of items this instance skipped because a producer overran it. Requires traits.enable_loss_detection, which is on by default; returns 0 when it is off. The counter is per-instance, not shared through the segment, and is cleared by reset().

initial_reading_index() -> int

Cursor for a late-joining consumer: 0 for a newly created queue, or the current writing index of a queue that was opened. Starting a reader here skips the backlog.

reset()

Clear the queue and rewind it to its initial state. Not thread-safe: call it only when no other thread or process is touching the queue. Readers holding a pre-reset() cursor recover only if they were built with enable_reset_check.

__getitem__(index) -> memoryview

Get memoryview for writing to reserved slot.

Parameters:

  • index (int): Index from reserve()

Returns:

  • memoryview: View into the data array

close()

Close the shared memory connection. Always call this before unlinking.

unlink()

Delete the shared memory segment. Only call from the process that created it.

AtomicCursor

The AtomicCursor class enables multi-consumer work-stealing patterns by providing an atomic read cursor that multiple consumers can coordinate through. Works in both local mode (multi-threading) and shared memory mode (multi-process).

__init__(buffer, offset=0)

Create an atomic cursor wrapper around a memory buffer.

Parameters:

  • buffer (memoryview or bytearray): Memory buffer
    • For local mode (threading): use bytearray(8)
    • For shared memory mode (multiprocess): use SharedMemory.buf
  • offset (int, optional): Byte offset in buffer (default 0)

Local Mode Example (Multi-Threading):

from slick_queue_py import AtomicCursor

# Create local cursor for multi-threading
cursor_buf = bytearray(8)
cursor = AtomicCursor(cursor_buf, 0)
cursor.store(0)  # Initialize to 0

Shared Memory Mode Example (Multi-Process):

from multiprocessing import shared_memory
from slick_queue_py import AtomicCursor

# Create shared cursor for multi-process
cursor_shm = shared_memory.SharedMemory(name='cursor', create=True, size=8)
cursor = AtomicCursor(cursor_shm.buf, 0)
cursor.store(0)  # Initialize to 0

load() -> int

Load the cursor value with atomic acquire semantics.

Returns:

  • int: Current cursor value

store(value)

Store a new cursor value with atomic release semantics.

Parameters:

  • value (int): New cursor value

compare_exchange_weak(expected, desired) -> Tuple[bool, int]

Atomically compare and swap the cursor value.

Parameters:

  • expected (int): Expected cursor value
  • desired (int): Desired cursor value

Returns:

  • Tuple[bool, int]: (success, actual_value)

Note: This is used internally by read(atomic_cursor) and typically doesn't need to be called directly.

Configuring Features (Traits)

Optional features are selected per queue through a traits argument, mirroring the Traits template parameter of C++ slick::queue<T, Traits>. Subclass QueueTraits and override only what you need:

from slick_queue_py import SlickQueue, QueueTraits

class MyTraits(QueueTraits):
    enable_reset_check = True   # opt in
    enable_read_last = False    # opt out

lean = SlickQueue(size=1024, element_size=8, traits=MyTraits)
standard = SlickQueue(size=1024, element_size=8)   # default traits - both can coexist
Trait Default Effect when enabled
enable_read_last True publish() maintains a last-published index so read_last() works. Costs one CAS per publish.
enable_reset_check False read() loads the producer's reservation counter and rewinds the cursor to 0 if it has run past it, which happens only when reset() rewound the counter. Without this, a reader holding a pre-reset() cursor returns None indefinitely and then skips the start of the new generation.
enable_loss_detection True Per-instance skipped-item counter, reported by loss_count().
enable_cpu_relax True Yield-based backoff on contended CAS loops.

There is one traits type and one default. C++ needs two (queue_traits and debug_queue_traits, selected by NDEBUG) so that a debug/release mismatch across translation units fails to link instead of silently violating the ODR - Python has no translation units, no ODR and no linker, and no debug/release build to key them off.

Loss detection is on by default here, unlike the C++ Release default. C++ pays a cacheline and an atomic fetch_add for it; in Python it measures at +0.2% on a reader that keeps up, and loss_count() is the only signal a consumer has that it was overrun - which also makes it the only way to know the bytes read() returned may have been overwritten in flight. __debug__ would have been the obvious analogue of NDEBUG, but it is about stripping asserts, and keying the counter to it would make loss_count() silently return 0 under python -O - hiding exactly the condition it exists to report. Subclass QueueTraits with enable_loss_detection = False for maximum throughput.

Notes:

  • read_last() requires enable_read_last. Calling it otherwise raises RuntimeError, not a silent fallback to the old reserved-cursor heuristic.
  • enable_read_last must match across a shared-memory segment. It is the one trait that changes the shared header protocol, so the creator records it in the segment's layout marker ('SLQ1' when the last-published index is maintained, 'SLQ0' when it is not) and every attacher checks it - including C++ peers, which use the same marker. A peer that disagrees is rejected with a RuntimeError at construction instead of silently corrupting the other side's view, in either direction: an attacher that expects the index would read a counter nobody writes, and one that does not maintain it would freeze read_last() for every peer that does. The other traits are local to each process and can differ freely on one segment.
  • A misspelled override is silent. enable_reset_chek = True in a subclass leaves the inherited attribute visible and keeps the base value. validate_traits() catches a wrong type, but cannot catch a typo.
  • Traits are snapshotted at construction. A traits type is an ordinary class, so its attributes stay writable, but the queue commits to the configuration once - it writes the layout marker from it and creates the optional atomics from it. Mutating the class afterwards therefore has no effect on queues already built from it, and q.traits is a read-only snapshot that always describes what that queue actually does. Build a new queue to change a setting.

Memory Layout

The queue uses the same memory layout as C++ slick::queue<T>:

Offset | Size          | Content
-------|---------------|------------------
0      | 8 bytes       | reserved_info (atomic uint64: 48-bit index, 16-bit size)
8      | 4 bytes       | uint32_t size (queue capacity)
12     | 4 bytes       | uint32_t element_size
16     | 8 bytes       | uint64_t last_published index (atomic)
24     | 4 bytes       | uint32_t header_magic - 'SLQ' + feature nibble
28     | 20 bytes      | padding (reserved)
48     | 4 bytes       | uint32_t init_state (atomic)
52     | 12 bytes      | padding (to 64 bytes)
64     | 16*size bytes | slot array
       | per slot:     |
       |   0-7         |   uint64_t data_index (atomic)
       |   8-11        |   uint32_t size (atomic)
       |   12-15       |   padding
64+... | elem*size     | data array

The header magic is the bytes 'SLQ' followed by an ASCII digit whose low nibble carries the shared-layout features the creator was built with:

Marker Value Meaning
'SLQ1' 0x534C5131 The last-published index at offset 16 is maintained
'SLQ0' 0x534C5130 It is not - read_last() is unavailable to every peer

Bits 1-3 of the nibble are reserved and must be 0; a marker that sets one is rejected as newer than this build understands. Segments created before slick-queue v1.4.0 carry no marker at all and are rejected at attach time.

Platform Support

Fully Supported (Lock-Free)

  • Windows x86-64: Uses C++ extension (atomic_ops_ext.pyd) with std::atomic
  • Linux x86-64: Uses C++ extension (atomic_ops_ext.so) with std::atomic, fallback to libatomic
  • macOS x86-64: Uses C++ extension (atomic_ops_ext.so) with std::atomic, fallback to compiler builtins

Platform-specific atomic operation implementations:

  • All platforms: The atomic_ops_ext C++ extension is now used on all platforms for the most reliable cross-process atomic operations
  • Fallback support: Linux/macOS can fall back to libatomic or compiler builtins if the extension isn't available

Building and Installation

The C++ extension is built automatically during installation:

# Install with automatic extension build
pip install -e .

# Or build manually first
python setup.py build_ext --inplace
pip install -e .

Build requirements:

  • Windows: Visual Studio 2017+ or MSVC build tools
  • Linux: GCC 5+ or Clang 3.8+
  • macOS: Xcode command line tools (clang)
  • All platforms: Python development headers (included with standard Python installation)

The extension will be built as:

  • Windows: atomic_ops_ext.cp3XX-win_amd64.pyd
  • Linux: atomic_ops_ext.cpython-3XX-x86_64-linux-gnu.so
  • macOS: atomic_ops_ext.cpython-3XX-darwin.so

(where XX is your Python version, e.g., 312 for Python 3.12)

Requirements for Lock-Free Operation

All platforms require hardware support for lock-free atomic operations:

  • x86-64 CPU with CMPXCHG16B instruction (Intel since ~2006, AMD since ~2007)
  • For C++/Python interoperability, both must use the same atomic hardware instructions
  • No fallback implementation exists - lock-free atomics are mandatory for multi-producer queues

Why no fallback? The queue requires true atomic CAS operations for correctness in multi-producer scenarios. A lock-based fallback would:

  • Break binary compatibility with C++ SlickQueue
  • Fail to work correctly in multi-process scenarios (Python ↔ C++)
  • Not provide the performance guarantees of a lock-free queue

Not Supported

  • 32-bit platforms (no 16-byte atomic CAS)
  • ARM64 (requires ARMv8.1+ CASP instruction - future support planned)
  • CPUs without CMPXCHG16B support (very old x86-64 CPUs from before 2006)

Check platform support:

from atomic_ops import check_platform_support

supported, message = check_platform_support()
print(f"Platform: {message}")

Performance

Typical throughput on modern hardware (x86-64):

  • Single producer/consumer: ~5-10M items/sec
  • 4 producers/1 consumer: ~3-8M items/sec
  • High contention (8+ producers): ~1-5M items/sec

Performance depends on:

  • CPU cache topology
  • Queue size (smaller = more contention)
  • Item size
  • Memory bandwidth

Advanced Usage

Batch Operations

Reserve and publish multiple elements at once:

# Reserve 10 elements
idx = q.reserve(10)

# Write data to each slot
for i in range(10):
    element = q[idx + i]
    element[:data_len] = data[i]

# Publish all 10 elements at once
q.publish(idx, 10)

Wrap-Around Handling

The queue automatically handles ring buffer wrap-around:

# Queue with size=8
q = SlickQueue(name='wrap_test', size=8, element_size=32)

# Reserve more items than queue size - wraps automatically
for i in range(100):
    idx = q.reserve()
    q[idx][:4] = struct.pack("<I", i)
    q.publish(idx)

Testing

Python Tests

Run the Python test suite:

# Atomic operations tests (clean output)
python tests/run_test.py tests/test_atomic_ops.py

# Basic queue tests (clean output)
python tests/run_test.py tests/test_queue.py

# Local mode tests
python tests/test_local_mode.py

# Multi-producer/consumer tests
# Note: If tests fail with "File exists" errors, run cleanup first:
python tests/cleanup_shm.py
python tests/test_multi_producer.py

# Traits, layout marker, and cross-peer feature mismatch
python tests/test_traits.py

# reset() recovery (enable_reset_check)
python tests/test_reset_detection.py

# Wrapping-producer record/size invariants
python tests/test_wrap_invariants.py

Or run everything with pytest:

python -m pytest tests/

C++/Python Interoperability Tests

Build and run comprehensive interop tests:

# 1. Build C++ test programs with CMake
mkdir build && cd build
cmake ..
cmake --build .

# 2. Run interoperability test suite
cd ..
python tests/test_interop.py

# Or run specific tests:
python tests/test_interop.py --test python_producer_cpp_consumer
python tests/test_interop.py --test cpp_producer_python_consumer
python tests/test_interop.py --test multi_producer_interop
python tests/test_interop.py --test stress_interop
python tests/test_interop.py --test cpp_shm_creation

The interop tests verify:

  • Python → C++: Python producers write data that C++ consumers read
  • C++ → Python: C++ producers write data that Python consumers read
  • Mixed Multi-Producer: Multiple C++ and Python producers writing to same queue
  • Stress Test: High-volume bidirectional communication
  • SHM created by C++: C++ producers create the SHM and write data that Python consumers read

Note on Windows: If child processes from previous test runs don't terminate properly, you may need to manually kill orphaned python.exe processes before running tests again.

Known Issues

  1. Buffer Cleanup Warning: You may see a BufferError: cannot close exported pointers exist warning during garbage collection. This is a harmless warning caused by Python's ctypes creating internal buffer references that persist beyond explicit cleanup. It occurs during program exit and does not affect functionality, performance, or correctness. The queue works perfectly despite this warning.

  2. UserWarning: On Linux you may see UserWarning: resource_tracker: There appear to be 4 leaked shared_memory objects to clean up at shutdown. This is a harmless warning caused by Python's ctypes creating internal buffer references that persist beyond explicit cleanup. It occurs during program exit and does not affect functionality, performance, or correctness. The queue works perfectly despite this warning.

Architecture

Atomic Operations

The queue uses platform-specific atomic operations:

  • 8-byte CAS: For reserved_info structure (multi-producer coordination)
  • 8-byte CAS: For slot data_index fields (publish/read synchronization)
  • Memory barriers: Acquire/release semantics for proper ordering

Memory Ordering

  • reserve(): Uses memory_order_release on successful CAS
  • publish(): Writes slot.size, then stores data_index with memory_order_release
  • read() / read_last(): Load data_index with memory_order_acquire, read slot.size once, then re-validate data_index before using either

This ensures:

  • All writes to data are visible before publishing
  • All reads of data happen after acquiring the index
  • The returned (data, size) pair always describes one and the same record, even when a wrapping producer recycles the slot mid-read - the re-validation is a seqlock bracket around the size load, and a reader that loses the race retries rather than returning a torn pair
  • No reordering that could cause data races

What is not guaranteed: the queue is lossy. A producer writes an element's data before it publishes the slot's new index, so a consumer that has been lapped can copy bytes the producer is midway through overwriting - nothing in the slot can detect this, because the index has not changed yet. C++ has the identical hazard and hands back a pointer whose target is the producer's to overwrite; Python copies, so the copy can hold a newer record's bytes under the older record's cursor. Size the queue so consumers keep up, and use loss_count() to detect when they have not.

Comparison with C++

Feature C++ Python
Multi-producer ✅ ✅
Multi-consumer (work-stealing) ✅ ✅ (with AtomicCursor)
Lock-free (x86-64) ✅ ✅
Memory layout Reference Matches exactly
Performance Baseline ~50-80% of C++
Ease of use Medium High
read(int) single-consumer ✅ ✅
read(atomic cursor) multi-consumer ✅ ✅
Feature traits Template parameter traits= argument
Shared layout marker 'SLQ1' / 'SLQ0' Same, and validated against C++ peers

Contributing

Issues and pull requests welcome at SlickQuant/slick-queue-py.

License

MIT License - see LICENSE file for details.

Made with ⚡ by SlickQuant

Release files for slick-queue-py 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for slick-queue-py 2.0.0
File Size Uploaded
slick_queue_py-2.0.0.tar.gz 96.0 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for slick-queue-py 2.0.0
File
slick_queue_py-2.0.0-cp313-cp313-win_amd64.whl CPython 3.13 CPython 3.13 Windows x86-64 Details
slick_queue_py-2.0.0-cp313-cp313-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.17+ x86-64, Linux glibc 2.5+ x86-64 Details
slick_queue_py-2.0.0-cp313-cp313-macosx_10_13_universal2.whl CPython 3.13 CPython 3.13 macOS 10.13+ universal2 (ARM64, x86-64) Details
slick_queue_py-2.0.0-cp312-cp312-win_amd64.whl CPython 3.12 CPython 3.12 Windows x86-64 Details
slick_queue_py-2.0.0-cp312-cp312-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.5+ x86-64, Linux glibc 2.17+ x86-64 Details
slick_queue_py-2.0.0-cp312-cp312-macosx_10_13_universal2.whl CPython 3.12 CPython 3.12 macOS 10.13+ universal2 (ARM64, x86-64) Details
slick_queue_py-2.0.0-cp311-cp311-win_amd64.whl CPython 3.11 CPython 3.11 Windows x86-64 Details
slick_queue_py-2.0.0-cp311-cp311-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.5+ x86-64, Linux glibc 2.17+ x86-64 Details
slick_queue_py-2.0.0-cp311-cp311-macosx_10_9_universal2.whl CPython 3.11 CPython 3.11 macOS 10.9+ universal2 (ARM64, x86-64) Details
slick_queue_py-2.0.0-cp310-cp310-win_amd64.whl CPython 3.10 CPython 3.10 Windows x86-64 Details
slick_queue_py-2.0.0-cp310-cp310-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.10 CPython 3.10 Linux glibc 2.5+ x86-64, Linux glibc 2.17+ x86-64 Details
slick_queue_py-2.0.0-cp310-cp310-macosx_10_9_universal2.whl CPython 3.10 CPython 3.10 macOS 10.9+ universal2 (ARM64, x86-64) Details
slick_queue_py-2.0.0-cp39-cp39-win_amd64.whl CPython 3.9 CPython 3.9 Windows x86-64 Details
slick_queue_py-2.0.0-cp39-cp39-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.9 CPython 3.9 Linux glibc 2.5+ x86-64, Linux glibc 2.17+ x86-64 Details
slick_queue_py-2.0.0-cp39-cp39-macosx_10_9_universal2.whl CPython 3.9 CPython 3.9 macOS 10.9+ universal2 (ARM64, x86-64) Details
slick_queue_py-2.0.0-cp38-cp38-win_amd64.whl CPython 3.8 CPython 3.8 Windows x86-64 Details
slick_queue_py-2.0.0-cp38-cp38-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.8 CPython 3.8 Linux glibc 2.5+ x86-64, Linux glibc 2.17+ x86-64 Details
slick_queue_py-2.0.0-cp38-cp38-macosx_10_9_universal2.whl CPython 3.8 CPython 3.8 macOS 10.9+ universal2 (ARM64, x86-64) Details

Total release size: 830.1 kB

Release files / slick_queue_py-2.0.0.tar.gz

Download URL slick_queue_py-2.0.0.tar.gz
Size 96.0 kB
Tags Source
SHA-256 checksum
How to use checksums
cd270b577db5c599366c43d46ca9bb8cef99a3bbfd17954cc7ee63994d020655
BLAKE2b-256 checksum
How to use checksums
6d519faf6cdc9df87bdb467f287660bbb43248903e42e8331019d0abbda74f4e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp313-cp313-win_amd64.whl

Download URL slick_queue_py-2.0.0-cp313-cp313-win_amd64.whl
Size 37.0 kB
Tags CPython 3.13 Windows x86-64
SHA-256 checksum
How to use checksums
6f944037fb88d78fc63f070353b40d9ab71ab5775bbb4525153e0f4a7f9eefff
BLAKE2b-256 checksum
How to use checksums
8d681743c98b35d2ea55b94545517854cb9aae165034db913b8c288d3e5e2b41
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp313-cp313-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL slick_queue_py-2.0.0-cp313-cp313-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 49.4 kB
Tags CPython 3.13 Linux glibc 2.17+ x86-64 Linux glibc 2.5+ x86-64
SHA-256 checksum
How to use checksums
6352b966714a266f0dcada648a7948e5dfb47f514347b697ef1d5e265683f131
BLAKE2b-256 checksum
How to use checksums
e2cb5b2c442cd04d6b8a8cea22680af3715545f8b946dd72ed8278281e8b58dc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp313-cp313-macosx_10_13_universal2.whl

Download URL slick_queue_py-2.0.0-cp313-cp313-macosx_10_13_universal2.whl
Size 35.9 kB
Tags CPython 3.13 macOS 10.13+ universal2 (ARM64, x86-64)
SHA-256 checksum
How to use checksums
aaa7b16e6e774b1711434c4bfe4854d50628d77507f9afa48c8e8719a06b919a
BLAKE2b-256 checksum
How to use checksums
a331f1c8d563b6886e7d83d4ca804485d695a9af2044088c4bf0fe93bee724e2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp312-cp312-win_amd64.whl

Download URL slick_queue_py-2.0.0-cp312-cp312-win_amd64.whl
Size 37.0 kB
Tags CPython 3.12 Windows x86-64
SHA-256 checksum
How to use checksums
f600f9a89b1661f89189bb48879ab7ed141bd05e03a14c33eea35a676bb7815e
BLAKE2b-256 checksum
How to use checksums
3be6490c0045e32cb75f8c65230be36e83ecf97c1404e0ff34d4ff3eb2fc4fd2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp312-cp312-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL slick_queue_py-2.0.0-cp312-cp312-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 49.4 kB
Tags CPython 3.12 Linux glibc 2.17+ x86-64 Linux glibc 2.5+ x86-64
SHA-256 checksum
How to use checksums
ac1f8529f8243747392fa64542b53201020b436ae95807e6e0c55ce8533bff53
BLAKE2b-256 checksum
How to use checksums
3f915225d016ca3bc6ba532c07745965a00ddf51b8563ffe45b86ec82b543ecb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp312-cp312-macosx_10_13_universal2.whl

Download URL slick_queue_py-2.0.0-cp312-cp312-macosx_10_13_universal2.whl
Size 35.9 kB
Tags CPython 3.12 macOS 10.13+ universal2 (ARM64, x86-64)
SHA-256 checksum
How to use checksums
f87795200c0e9b4b11533d27b94c1de30ac330cad5e3f58065b37b648ab0ff72
BLAKE2b-256 checksum
How to use checksums
ee403ab55550f1a9ec5c19174e7b131bcbb167d308201cde6fed87a6b3aafd96
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp311-cp311-win_amd64.whl

Download URL slick_queue_py-2.0.0-cp311-cp311-win_amd64.whl
Size 37.0 kB
Tags CPython 3.11 Windows x86-64
SHA-256 checksum
How to use checksums
45b56c149e8ab055e6f7b5f86fdfd1d403455c73f7898a9ee77ceaa6e108128c
BLAKE2b-256 checksum
How to use checksums
54d38a698aa6fdca7cf4820fc7921d64a47b79b809efdfabf847dc72fa51a25f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp311-cp311-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL slick_queue_py-2.0.0-cp311-cp311-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 49.5 kB
Tags CPython 3.11 Linux glibc 2.17+ x86-64 Linux glibc 2.5+ x86-64
SHA-256 checksum
How to use checksums
bf8f26c75033779b84390415081687883891badc6598059a17b5d2e02ad5edd8
BLAKE2b-256 checksum
How to use checksums
a799e4e803fedde2bcbdedcf5eb9d7fd70d30a9d6f471f4d0cd32abfde32942a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp311-cp311-macosx_10_9_universal2.whl

Download URL slick_queue_py-2.0.0-cp311-cp311-macosx_10_9_universal2.whl
Size 35.9 kB
Tags CPython 3.11 macOS 10.9+ universal2 (ARM64, x86-64)
SHA-256 checksum
How to use checksums
41e0612d4663391d79239010948db3c1c6f614987fba09a4278000eb8f8c9edc
BLAKE2b-256 checksum
How to use checksums
2361c3b887bd8120d42811bfc004377924fa9677d2980f4b522954885357dadf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp310-cp310-win_amd64.whl

Download URL slick_queue_py-2.0.0-cp310-cp310-win_amd64.whl
Size 37.0 kB
Tags CPython 3.10 Windows x86-64
SHA-256 checksum
How to use checksums
0895d045f66929032d581db95ab759b1b1084ccc26530d026e3f0924c48b0336
BLAKE2b-256 checksum
How to use checksums
19f2ec00e5c5379a2153dd210ed8515e433933f315abe3a0bd816dcba604d25f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp310-cp310-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL slick_queue_py-2.0.0-cp310-cp310-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 49.4 kB
Tags CPython 3.10 Linux glibc 2.17+ x86-64 Linux glibc 2.5+ x86-64
SHA-256 checksum
How to use checksums
2a44061d5e8357990407b6d44b96add5d38d59f47a86010754ac662b06493e12
BLAKE2b-256 checksum
How to use checksums
f5cd3390d0e8585d4f326314b66ee62b548fd994e1c76869298da52073c2666d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp310-cp310-macosx_10_9_universal2.whl

Download URL slick_queue_py-2.0.0-cp310-cp310-macosx_10_9_universal2.whl
Size 35.9 kB
Tags CPython 3.10 macOS 10.9+ universal2 (ARM64, x86-64)
SHA-256 checksum
How to use checksums
31c651b4561697cddfba8eec14f4010ebfadfd7dd77d485f35f83fccc6c96e2b
BLAKE2b-256 checksum
How to use checksums
64569c80294085f6b889e638ca82aeffadd2d438a41191c89207355c2e307afa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp39-cp39-win_amd64.whl

Download URL slick_queue_py-2.0.0-cp39-cp39-win_amd64.whl
Size 37.0 kB
Tags CPython 3.9 Windows x86-64
SHA-256 checksum
How to use checksums
4d13ea702288ae36d21db4ee02d430a2add8509df5c948e19cbffae8d8a62e60
BLAKE2b-256 checksum
How to use checksums
ab61bef943d2819e197a36a3a74459e6d61f4770ef8b8760a3af5a98e25bd13d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp39-cp39-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL slick_queue_py-2.0.0-cp39-cp39-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 49.2 kB
Tags CPython 3.9 Linux glibc 2.17+ x86-64 Linux glibc 2.5+ x86-64
SHA-256 checksum
How to use checksums
799bd3906615fc293dfdf503ba934febc8a4ac508c41f7e96edd4bc70d0f0a5a
BLAKE2b-256 checksum
How to use checksums
72359f8c62fca677ba862d80dadf6ece34ea6001e4f17cd321c37c5daa2d7ea4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp39-cp39-macosx_10_9_universal2.whl

Download URL slick_queue_py-2.0.0-cp39-cp39-macosx_10_9_universal2.whl
Size 35.9 kB
Tags CPython 3.9 macOS 10.9+ universal2 (ARM64, x86-64)
SHA-256 checksum
How to use checksums
e44e02e935dc549f67c584265c35f102c0b80d884c7c08015d0f62b5d4b91124
BLAKE2b-256 checksum
How to use checksums
4af5456a5e0ae8fce5f583a06d21150a85d1963c33b54d9401f21f4a5a32ca29
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp38-cp38-win_amd64.whl

Download URL slick_queue_py-2.0.0-cp38-cp38-win_amd64.whl
Size 37.0 kB
Tags CPython 3.8 Windows x86-64
SHA-256 checksum
How to use checksums
1018d04268fe523f1ba2e38227f1d702eb8d01c1a18ab7cbc84e8b8d57588a43
BLAKE2b-256 checksum
How to use checksums
84b22f71f32ce8d8231ce74919d74c777eccfb69d476043bef2323e7e793fc59
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp38-cp38-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL slick_queue_py-2.0.0-cp38-cp38-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 49.6 kB
Tags CPython 3.8 Linux glibc 2.17+ x86-64 Linux glibc 2.5+ x86-64
SHA-256 checksum
How to use checksums
ef61f6fa9b36e502e5e64cc3d16becc750be8dcf3ffbe7061e23dcd83b4aeb86
BLAKE2b-256 checksum
How to use checksums
470dbe666aa2e35faf1ff141a801067439ec7f01e7637a90cdffaba72fac1f0e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / slick_queue_py-2.0.0-cp38-cp38-macosx_10_9_universal2.whl

Download URL slick_queue_py-2.0.0-cp38-cp38-macosx_10_9_universal2.whl
Size 35.7 kB
Tags CPython 3.8 macOS 10.9+ universal2 (ARM64, x86-64)
SHA-256 checksum
How to use checksums
0ce579ced0abc85872cb97c1e469fbf84ebc7933670efe34be45045a61f3f629
BLAKE2b-256 checksum
How to use checksums
9f4bda18cb516c6ffea8791db85e8e30d33a743b9ca90453cb52831df94580a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

2.1.0

19 release files

This release

2.0.0 This release

19 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page