Skip to main content

PegaFlow Python Package

High-performance key-value storage engine with Python bindings, built with Rust and PyO3.

Features

  • PegaEngine: Fast Rust-based key-value storage with Python bindings
  • PegaKVConnector: vLLM KV connector for distributed inference with KV cache transfer

Installation

From Source

# Install maturin if you haven't already
pip install maturin

# Build and install in development mode
cd python
maturin develop

# Or build a wheel
maturin build --release

From PyPI (coming soon)

pip install pegaflow

Usage

Basic KV Storage

from pegaflow import PegaEngine

# Create a new engine
engine = PegaEngine()

# Store key-value pairs
engine.put("name", "PegaFlow")
engine.put("version", "0.1.0")

# Retrieve values
name = engine.get("name")  # Returns "PegaFlow"
missing = engine.get("nonexistent")  # Returns None

# Remove keys
removed = engine.remove("name")  # Returns "PegaFlow"

vLLM KV Connector

from vllm import LLM
from vllm.distributed.kv_transfer.kv_transfer_agent import KVTransferConfig

# Configure vLLM to use PegaKVConnector
kv_transfer_config = KVTransferConfig(
    kv_connector="PegaKVConnector",
    kv_role="kv_both",
    kv_connector_module_path="pegaflow.connector",
)

# Create LLM with KV transfer enabled
llm = LLM(
    model="gpt2",
    kv_transfer_config=kv_transfer_config,
)

Connector Modes

PegaKVConnector defaults to read_write: it queries PegaFlow for reusable KV blocks, loads matched blocks into vLLM, and saves newly computed full blocks back to PegaFlow.

Set pegaflow.mode to save_only when another vLLM connector is responsible for reads and PegaFlow should only persist KV blocks for later reuse. This is intended for MultiConnector decode-side setups where an upstream connector owns the external hit/load path, while PegaFlow records the resulting KV cache. In save_only mode, PegaFlow does not query or load KV blocks.

vllm serve Qwen/Qwen3-0.6B \
  --kv-transfer-config '{
    "kv_connector": "MultiConnector",
    "kv_role": "kv_both",
    "kv_connector_extra_config": {
      "connectors": [
        {
          "kv_connector": "<external-read-connector>",
          "kv_role": "kv_both"
        },
        {
          "kv_connector": "PegaKVConnector",
          "kv_role": "kv_both",
          "kv_connector_module_path": "pegaflow.connector",
          "kv_connector_extra_config": {
            "pegaflow.mode": "save_only"
          }
        }
      ]
    }
  }'

Valid values are read_write and save_only.

TP Shards Across Hosts

CUDA IPC is host-local. When one tensor-parallel replica spans multiple hosts, run one PegaFlow server on each host and configure the connector with every server endpoint in global TP-rank order:

{
  "kv_connector": "PegaKVConnector",
  "kv_role": "kv_both",
  "kv_connector_module_path": "pegaflow.connector",
  "kv_connector_extra_config": {
    "pegaflow.tp_shard_endpoints": [
      "http://host-a:50055",
      "http://host-b:50055"
    ]
  }
}

For TP8 and two endpoints, global ranks 0-3 register with the first server and ranks 4-7 register with the second. Each server sees a local TP4 topology and must manage the four GPUs on its own host. Every vLLM process must receive the same ordered endpoint list.

The scheduler queries every shard and only reuses the prefix available from all of them. Each worker loads with the lease issued by its local server. The connector gives every shard a distinct namespace, so deployments with a different host split cannot reuse an incompatible cache layout.

TP sharding currently requires equal contiguous shards and TP-only parallelism. Pipeline, decode-context, and prefill-context parallelism are rejected when more than one endpoint is configured.

P/D Partial Tail Blocks

vLLM normally exposes hashes only for complete KV blocks. In a P/D deployment, enable pegaflow.pd_tail_save on prefill and pegaflow.pd_tail_load on decode to reuse the final partial prompt block as well. Start both vLLM processes with the same explicit PYTHONHASHSEED and --prefix-caching-hash-algo xxhash_cbor.

Prefill: {"pegaflow.pd_tail_save": true}

Decode: {"pegaflow.pd_tail_load": true, "pegaflow.wait_for_full_prefix": true}

pegaflow.wait_for_full_prefix makes decode wait (up to 30s) until the full prompt prefix is fetchable from a remote node via MetaServer + RDMA. It only applies when prefill and decode run separate engines; it does not observe saves landing in a shared/local engine and has no effect when RDMA is not configured.

Development

See the examples directory for more usage examples.

Testing

Running Unit Tests

The test suite includes integration tests that verify the EngineRpcClient can correctly communicate with a running pegaflow-server instance.

Prerequisites

  1. Build the Rust extension:

    cd python
    maturin develop --release
    
  2. Build the server binary:

    cd ..
    cargo build --release --bin pegaflow-server
    
  3. Ensure CUDA is available (tests require GPU):

    python -c "import torch; assert torch.cuda.is_available()"
    

Running Tests

cd python

# Run all tests
pytest tests/ -v

# Run specific test file
pytest tests/test_engine_client.py -v

# Run with coverage
pytest tests/ --cov=pegaflow --cov-report=html

Test Structure

  • tests/conftest.py: Contains pytest fixtures for:

    • pega_server: Automatically starts/stops pegaflow-server for integration tests
    • engine_client: Creates an EngineRpcClient connected to the test server
    • client_context: Provides a ClientContext representing a vLLM instance with GPU KV cache tensors
    • registered_instance: Provides a registered instance ID for query tests
  • tests/test_engine_client.py: Integration tests for:

    • Server connectivity
    • Query operations with various inputs

Test Fixtures

The ClientContext class abstracts a vLLM instance and provides:

  • register_kv_caches(): Register GPU KV cache tensors with the server
  • query(block_hashes): Query available blocks
  • unregister_context(): Unregister context from server

Example test usage:

def test_query(client_context):
    """Test query operation."""
    result = client_context.query([])
    assert result is not None

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pegaflow_llm-0.23.11-cp314-cp314-manylinux_2_34_x86_64.whl (8.0 MB view details)

Uploaded CPython 3.14manylinux: glibc 2.34+ x86-64

pegaflow_llm-0.23.11-cp313-cp313-manylinux_2_34_x86_64.whl (8.0 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.34+ x86-64

pegaflow_llm-0.23.11-cp312-cp312-manylinux_2_34_x86_64.whl (8.0 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.34+ x86-64

pegaflow_llm-0.23.11-cp311-cp311-manylinux_2_34_x86_64.whl (8.0 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.34+ x86-64

pegaflow_llm-0.23.11-cp310-cp310-manylinux_2_34_x86_64.whl (8.0 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.34+ x86-64

File details

Details for the file pegaflow_llm-0.23.11-cp314-cp314-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm-0.23.11-cp314-cp314-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 a0aaaf6dee17d8936f4e56fbfb6ca09656ed15acd9e6a4e51f5b04614065369a
MD5 002b5190e193e6326fc7cf957475bc51
BLAKE2b-256 8b15fd7a9f3d1781342b62b4731100ebf5fae93039f7b90f37bbf7dd7289f8fe

See more details on using hashes here.

File details

Details for the file pegaflow_llm-0.23.11-cp313-cp313-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm-0.23.11-cp313-cp313-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 5507840a1e3811a9e963c7937f69b719cde0f32c8c71ebfe79bec65d4c094617
MD5 18e13f9c473c4f14ae4604bb9c9d9ad4
BLAKE2b-256 3633fc0f951613a5b4cda5c4e0a0db877c114a4c81bc0cd188d1c7aafdc9e520

See more details on using hashes here.

File details

Details for the file pegaflow_llm-0.23.11-cp312-cp312-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm-0.23.11-cp312-cp312-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 ff03fff6f4705ace1c8a553a5743bb6a7e1937fad66f86021730f6eadc1d30de
MD5 4abf0d4d6239f4dbe2e8d169d09ecf43
BLAKE2b-256 6cb4086633eceb7a90a8b2de4cdb671fd7f83909337243ad510e3fdac1f997c4

See more details on using hashes here.

File details

Details for the file pegaflow_llm-0.23.11-cp311-cp311-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm-0.23.11-cp311-cp311-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 7fe279a280efca4251f2e7aa568f704375e646a622a1fbfbd6ec511a3f1e6388
MD5 5ba909097df760d287821c8ae5d1001e
BLAKE2b-256 e6a77c4521703d7f697cbc57da9962bf3f2365230a336db441a0275fccd14136

See more details on using hashes here.

File details

Details for the file pegaflow_llm-0.23.11-cp310-cp310-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for pegaflow_llm-0.23.11-cp310-cp310-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 99ad91b45a052b93042a7e98ea27411b153f1c376f6d9a79dc8d1eb3e684af00
MD5 ebcfea8a39f49c22c3f481e0b5fe8770
BLAKE2b-256 f0da597afa86fa36884c45d1f16dad70ad9beeb567478150ef54166d56528c2e

See more details on using hashes here.

Release history Release notifications | RSS feed

0.24.1

5 files

0.24.0

5 files

0.23.13

5 files

0.23.12

5 files

This release

0.23.11 This release

5 files

0.23.10

5 files

0.23.9

5 files

0.23.8

5 files

0.23.7

5 files

0.23.6

5 files

0.23.5

5 files

0.23.4

5 files

0.23.3

5 files

0.23.2

5 files

0.23.1

5 files

0.23.0

5 files

0.22.10

5 files

0.22.9

5 files

0.22.8

5 files

0.22.7

5 files

0.22.6

5 files

0.22.5

5 files

0.22.4

5 files

0.22.3

5 files

0.22.2

5 files

0.22.1

5 files

0.22.0

5 files

0.21.2

5 files

0.21.1

4 files

0.21.0

4 files

0.20.0

4 files

0.19.1

4 files

0.19.0

4 files

0.18.0

4 files

0.17.0

4 files

0.0.16

4 files

0.0.14

3 files

0.0.13

3 files

0.0.12

3 files

0.0.11

3 files

0.0.10

3 files

0.0.9

3 files

0.0.8

3 files

0.0.7

3 files

0.0.6

3 files

0.0.5

3 files

0.0.4

3 files

0.0.3

3 files

0.0.2

3 files

0.0.1

3 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page