Skip to main content

PegaFlow Python Package

High-performance key-value storage engine with Python bindings, built with Rust and PyO3.

Features

  • PegaEngine: Fast Rust-based key-value storage with Python bindings
  • PegaKVConnector: vLLM KV connector for distributed inference with KV cache transfer

Installation

From Source

# Install maturin if you haven't already
pip install maturin

# Build and install in development mode
cd python
maturin develop

# Or build a wheel
maturin build --release

From PyPI (coming soon)

pip install pegaflow

Usage

Basic KV Storage

from pegaflow import PegaEngine

# Create a new engine
engine = PegaEngine()

# Store key-value pairs
engine.put("name", "PegaFlow")
engine.put("version", "0.1.0")

# Retrieve values
name = engine.get("name")  # Returns "PegaFlow"
missing = engine.get("nonexistent")  # Returns None

# Remove keys
removed = engine.remove("name")  # Returns "PegaFlow"

vLLM KV Connector

from vllm import LLM
from vllm.distributed.kv_transfer.kv_transfer_agent import KVTransferConfig

# Configure vLLM to use PegaKVConnector
kv_transfer_config = KVTransferConfig(
    kv_connector="PegaKVConnector",
    kv_role="kv_both",
    kv_connector_module_path="pegaflow.connector",
)

# Create LLM with KV transfer enabled
llm = LLM(
    model="gpt2",
    kv_transfer_config=kv_transfer_config,
)

Connector Modes

PegaKVConnector defaults to read_write: it queries PegaFlow for reusable KV blocks, loads matched blocks into vLLM, and saves newly computed full blocks back to PegaFlow.

Set pegaflow.mode to save_only when another vLLM connector is responsible for reads and PegaFlow should only persist KV blocks for later reuse. This is intended for MultiConnector decode-side setups where an upstream connector owns the external hit/load path, while PegaFlow records the resulting KV cache. In save_only mode, PegaFlow does not query or load KV blocks.

vllm serve Qwen/Qwen3-0.6B \
  --kv-transfer-config '{
    "kv_connector": "MultiConnector",
    "kv_role": "kv_both",
    "kv_connector_extra_config": {
      "connectors": [
        {
          "kv_connector": "<external-read-connector>",
          "kv_role": "kv_both"
        },
        {
          "kv_connector": "PegaKVConnector",
          "kv_role": "kv_both",
          "kv_connector_module_path": "pegaflow.connector",
          "kv_connector_extra_config": {
            "pegaflow.mode": "save_only"
          }
        }
      ]
    }
  }'

Valid values are read_write and save_only.

P/D Partial Tail Blocks

vLLM normally exposes hashes only for complete KV blocks. In a P/D deployment, enable pegaflow.pd_tail_save on prefill and pegaflow.pd_tail_load on decode to reuse the final partial prompt block as well. Start both vLLM processes with the same explicit PYTHONHASHSEED and --prefix-caching-hash-algo xxhash_cbor.

Prefill: {"pegaflow.pd_tail_save": true}

Decode: {"pegaflow.pd_tail_load": true, "pegaflow.wait_for_full_prefix": true}

pegaflow.wait_for_full_prefix makes decode wait (up to 30s) until the full prompt prefix is fetchable from a remote node via MetaServer + RDMA. It only applies when prefill and decode run separate engines; it does not observe saves landing in a shared/local engine and has no effect when RDMA is not configured.

Development

See the examples directory for more usage examples.

Testing

Running Unit Tests

The test suite includes integration tests that verify the EngineRpcClient can correctly communicate with a running pegaflow-server instance.

Prerequisites

  1. Build the Rust extension:

    cd python
    maturin develop --release
    
  2. Build the server binary:

    cd ..
    cargo build --release --bin pegaflow-server
    
  3. Ensure CUDA is available (tests require GPU):

    python -c "import torch; assert torch.cuda.is_available()"
    

Running Tests

cd python

# Run all tests
pytest tests/ -v

# Run specific test file
pytest tests/test_engine_client.py -v

# Run with coverage
pytest tests/ --cov=pegaflow --cov-report=html

Test Structure

  • tests/conftest.py: Contains pytest fixtures for:

    • pega_server: Automatically starts/stops pegaflow-server for integration tests
    • engine_client: Creates an EngineRpcClient connected to the test server
    • client_context: Provides a ClientContext representing a vLLM instance with GPU KV cache tensors
    • registered_instance: Provides a registered instance ID for query tests
  • tests/test_engine_client.py: Integration tests for:

    • Server connectivity
    • Query operations with various inputs

Test Fixtures

The ClientContext class abstracts a vLLM instance and provides:

  • register_kv_caches(): Register GPU KV cache tensors with the server
  • query(block_hashes): Query available blocks
  • unregister_context(): Unregister context from server

Example test usage:

def test_query(client_context):
    """Test query operation."""
    result = client_context.query([])
    assert result is not None

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pegaflow_llm_cu13-0.23.5-cp314-cp314-manylinux_2_34_x86_64.whl (7.9 MB view details)

Uploaded CPython 3.14manylinux: glibc 2.34+ x86-64

pegaflow_llm_cu13-0.23.5-cp313-cp313-manylinux_2_34_x86_64.whl (7.9 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.34+ x86-64

pegaflow_llm_cu13-0.23.5-cp312-cp312-manylinux_2_34_x86_64.whl (7.9 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.34+ x86-64

pegaflow_llm_cu13-0.23.5-cp311-cp311-manylinux_2_34_x86_64.whl (7.9 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.34+ x86-64

pegaflow_llm_cu13-0.23.5-cp310-cp310-manylinux_2_34_x86_64.whl (7.9 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.34+ x86-64

File details

Details for the file pegaflow_llm_cu13-0.23.5-cp314-cp314-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm_cu13-0.23.5-cp314-cp314-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 fe24abec04fa3228f592c2a3117ac0f382810bc62857591808f5d6e0708984e8
MD5 dcaf67f4b5d34fe0e191b1ecf16d456e
BLAKE2b-256 e3e4a6a281b565f40cd93c502616866e085690ca5a04b3f921f1caecc6d0220b

See more details on using hashes here.

File details

Details for the file pegaflow_llm_cu13-0.23.5-cp313-cp313-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm_cu13-0.23.5-cp313-cp313-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 329e58fbfea39f3d9c237ab4a8350c51e0099ec0e01a25cab9be0963271eae1c
MD5 76dd643a4c7de23ef1fb6709af292e9d
BLAKE2b-256 7f074a8fe10b34dd41ccdd946ed7db4e0fff399a7dea88ba6d29fa49abcc889a

See more details on using hashes here.

File details

Details for the file pegaflow_llm_cu13-0.23.5-cp312-cp312-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm_cu13-0.23.5-cp312-cp312-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 fbf5a7fa53b3782e95db2ca4c6682898ba5115b63399b4c24f792345d7b39629
MD5 9c6c8771a80c32d21357a07ad4dfc85e
BLAKE2b-256 bb8dc17f973a3e6bf5af4dddb48138c7f344b7ef39b65936817fcb403531fe3b

See more details on using hashes here.

File details

Details for the file pegaflow_llm_cu13-0.23.5-cp311-cp311-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm_cu13-0.23.5-cp311-cp311-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 419c228349d6229e1570e5f5bacc30677f49857c92aebd55864bd90748f72618
MD5 9fc0f39526ae560c677dbaedcae7f1af
BLAKE2b-256 ec080cb782af1b946dfb35ed42d8028552f75e7b2349dcaf0627a30d0859bc28

See more details on using hashes here.

File details

Details for the file pegaflow_llm_cu13-0.23.5-cp310-cp310-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm_cu13-0.23.5-cp310-cp310-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 39e6086ec4b97605999be12e40f5cd90e02e28dc8ff1ff9b6b9cf5b1a495acc9
MD5 8e713a9705fdcc995c48b867d62ebbd4
BLAKE2b-256 95174d1fb595109b0b57b7344b96fffe8688e4aff4a0ec2a0adfdf73f8ebc6cf

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page