Skip to main content

SemBlend vLLM Connector

CI License Python

vLLM KVConnector for SemBlend-backed semantic KV donor discovery.

This repo is the open-source adapter layer between vLLM and SemBlend.

SemBlend is a semantic KV reuse research library. It exists to evaluate when similar prompts may safely reuse or blend previously computed KV state. This connector exposes that work through vLLM's KVConnectorBase_V1 lifecycle.

Status

Experimental, with safe defaults. Verified paraphrase reuse in semantic_span_experimental mode has been validated end to end on stock vLLM 0.26 (Qwen2.5-7B, A10G) and through an llm-d gateway with semantic-affinity placement; start with docs/QUICKSTART_VLLM.md.

Default behavior is discovery-only:

  • exact vLLM prefix caching remains authoritative;
  • semantic lookup runs only after exact prefix coverage is insufficient;
  • the connector records donor hits, misses, and rejection reasons;
  • it returns (0, False) from get_num_new_matched_tokens() unless a configured materialization mode can prove a block-aligned exact token prefix or an explicitly opted-in isolated proof path;
  • normal vLLM execution continues on every provider error or unsupported case.

Install

From PyPI:

pip install "semblend-vllm-connector[semblend]"

Development:

pip install -e ".[semblend,dev]"

Run local checks:

make check

vLLM Configuration

Discovery-only mode:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-prefix-caching \
  --kv-transfer-config '{
    "kv_connector": "SemBlendVllmConnector",
    "kv_connector_module_path": "semblend_vllm_connector.connector",
    "kv_role": "kv_both",
    "kv_load_failure_policy": "recompute",
    "kv_connector_extra_config": {
      "mode": "discovery_only",
      "provider": "local",
      "min_prompt_tokens": 512,
      "min_similarity": 0.70
    }
  }'

SemBlend provider mode:

{
  "kv_connector": "SemBlendVllmConnector",
  "kv_connector_module_path": "semblend_vllm_connector.connector",
  "kv_role": "kv_both",
  "kv_load_failure_policy": "recompute",
  "kv_connector_extra_config": {
    "mode": "discovery_only",
    "provider": "semblend",
    "min_prompt_tokens": 512,
    "min_similarity": 0.70,
    "min_reuse_ratio": 0.50,
    "embedder_type": "minilm",
    "model_id": "meta-llama/Llama-3.1-8B-Instruct"
  }
}

Equivalent JSON examples live in examples/.

Optional audit stream for reproducible validation:

{
  "kv_connector_extra_config": {
    "audit_path": "/tmp/semblend-vllm-audit.jsonl",
    "log_decisions": true
  }
}

The audit file is JSONL. Runtime KV reuse should be counted only from runtime_materialized events. Semantic lookup hits and advertised loads are reported separately so benchmark runners can distinguish discovery from backend-confirmed materialization.

Modes

Mode Positive matched tokens? Purpose
discovery_only No Safe telemetry and workload qualification.
exact_prefix Only with engine-valid exact block refs Future safe materialization path.
request_only_experimental Yes, exact-token-prefix blocks by default Isolated validation mode; run with vLLM prefix caching disabled.
segmented_experimental Not enabled in this repo yet Requires segmented/sparse execution and recompute-boundary support.
semantic_span_experimental Yes, block-aligned donor spans with RoPE re-rotation Semantic reuse of non-identical prompts. Two lanes: verified paraphrase (whole-span serve gated by a fail-closed fact check; works on stock vLLM 0.26) and interior span (same content under a different wrapper; needs the scheduler re-consult patch in WorldFlowAI/vllm). See docs/QUICKSTART_VLLM.md.

In semantic_span_experimental the connector never advertises more than the donor KV it actually captured: spans trim to the stored donor window, a load that materializes zero layers fails loudly rather than decoding over uninitialized blocks, and every load leaves separate advertised and runtime_materialized audit records so reuse is only ever counted from the latter.

request_only_experimental defaults to exact-token-prefix materialization. The old zero-exact semantic proof behavior requires allow_non_identical_request_only=true or SEMBLEND_VLLM_ALLOW_NON_IDENTICAL_REQUEST_ONLY=1. Keep that flag limited to quality-gated validation experiments; it is not a production-safe substitute for a segmented/recompute engine path.

Safety Rules

The connector must not:

  • weaken exact prefix-cache semantics;
  • report semantic hits as computed tokens unless KV can actually be loaded;
  • publish non-identical semantic donor KV into vLLM's exact prefix cache;
  • treat non-identical semantic discovery as materializable unless an explicit validation flag is set and the run has separate quality gates;
  • cross model, tokenizer, adapter, or cache-salt namespaces;
  • fail inference because semantic lookup failed.

Repository Layout

src/semblend_vllm_connector/
  connector.py        vLLM KVConnectorBase_V1 implementation
  config.py           config/env parsing
  provider.py         provider protocol + local deterministic provider
  providers/
    semblend.py       lazy SemBlendPipeline adapter
  types.py            shared dataclasses/enums
  namespace.py        vLLM request namespace extraction

docs/
  ARCHITECTURE.md     detailed architecture and rollout plan
  SEMBLEND_PROVIDER.md
  VLLM_CONNECTOR_CONTRACT.md

examples/
  discovery_kv_transfer_config.json
  semblend_discovery_kv_transfer_config.json

Open Source Posture

This project follows the dynamic connector pattern used by mature vLLM KV cache projects: vLLM loads the connector from a Python module path, connector-specific settings live in kv_connector_extra_config, and unsafe materialization cases fail closed to normal vLLM prefill.

See:

Metadata

Release files for semblend-vllm-connector 0.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for semblend-vllm-connector 0.2.2
File Size Uploaded
semblend_vllm_connector-0.2.2.tar.gz 88.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for semblend-vllm-connector 0.2.2
File Interpreter ABI Platform
semblend_vllm_connector-0.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 132.3 kB

Release files / semblend_vllm_connector-0.2.2.tar.gz

Download URL semblend_vllm_connector-0.2.2.tar.gz
Size 88.2 kB
Tags Source
SHA-256 checksum
How to use checksums
d00f5da9321d0c49e696327c09647339fbb9b329c4bd8065316d84e678c9997d
BLAKE2b-256 checksum
How to use checksums
690fa8449e8642cc500f2f0d1e4dc17642fe7e29a3e97af7668cbb74ba4a2e32
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release files / semblend_vllm_connector-0.2.2-py3-none-any.whl

Download URL semblend_vllm_connector-0.2.2-py3-none-any.whl
Size 44.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d85adcf0aeec20d61890ad0ea28ba54aa657cd0eedccff85380065ae8a80fda0
BLAKE2b-256 checksum
How to use checksums
7c5efb491c3e6effa905b5daba21fcbe10702b8649e7e0f2d64ed8daeb68d206
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release history Release notifications | RSS feed

0.2.8

2 release files

This release

0.2.2 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page