Skip to main content

SemBlend vLLM Connector

CI License Python

vLLM KVConnector for SemBlend-backed semantic KV donor discovery.

This repo is the open-source adapter layer between vLLM and SemBlend.

SemBlend is a semantic KV reuse research library. It exists to evaluate when similar prompts may safely reuse or blend previously computed KV state. This connector exposes that work through vLLM's KVConnectorBase_V1 lifecycle.

Status

Experimental, with safe defaults. Verified paraphrase reuse in semantic_span_experimental mode has been validated end to end on stock vLLM 0.26 (Qwen2.5-7B, A10G) and through an llm-d gateway with semantic-affinity placement; start with docs/QUICKSTART_VLLM.md.

Default behavior is discovery-only:

  • exact vLLM prefix caching remains authoritative;
  • semantic lookup runs only after exact prefix coverage is insufficient;
  • the connector records donor hits, misses, and rejection reasons;
  • it returns (0, False) from get_num_new_matched_tokens() unless a configured materialization mode can prove a block-aligned exact token prefix or an explicitly opted-in isolated proof path;
  • normal vLLM execution continues on every provider error or unsupported case.

Install

From PyPI:

pip install "semblend-vllm-connector[semblend]"

Development:

pip install -e ".[semblend,dev]"

Run local checks:

make check

vLLM Configuration

Discovery-only mode:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-prefix-caching \
  --kv-transfer-config '{
    "kv_connector": "SemBlendVllmConnector",
    "kv_connector_module_path": "semblend_vllm_connector.connector",
    "kv_role": "kv_both",
    "kv_load_failure_policy": "recompute",
    "kv_connector_extra_config": {
      "mode": "discovery_only",
      "provider": "local",
      "min_prompt_tokens": 256,
      "min_similarity": 0.70
    }
  }'

SemBlend provider mode:

{
  "kv_connector": "SemBlendVllmConnector",
  "kv_connector_module_path": "semblend_vllm_connector.connector",
  "kv_role": "kv_both",
  "kv_load_failure_policy": "recompute",
  "kv_connector_extra_config": {
    "mode": "discovery_only",
    "provider": "semblend",
    "min_prompt_tokens": 256,
    "min_similarity": 0.70,
    "min_reuse_ratio": 0.50,
    "embedder_type": "minilm",
    "model_id": "meta-llama/Llama-3.1-8B-Instruct"
  }
}

Equivalent JSON examples live in examples/.

Optional audit stream for reproducible validation:

{
  "kv_connector_extra_config": {
    "audit_path": "/tmp/semblend-vllm-audit.jsonl",
    "log_decisions": true
  }
}

The audit file is JSONL. Runtime KV reuse should be counted only from runtime_materialized events. Semantic lookup hits and advertised loads are reported separately so benchmark runners can distinguish discovery from backend-confirmed materialization.

Modes

Mode Positive matched tokens? Purpose
discovery_only No Safe telemetry and workload qualification.
exact_prefix Only with engine-valid exact block refs Future safe materialization path.
request_only_experimental Yes, exact-token-prefix blocks by default Isolated validation mode; run with vLLM prefix caching disabled.
segmented_experimental Not enabled in this repo yet Requires segmented/sparse execution and recompute-boundary support.
semantic_span_experimental Yes, block-aligned donor spans with RoPE re-rotation Semantic reuse of non-identical prompts. Two lanes: verified paraphrase (whole-span serve gated by a fail-closed fact check; works on stock vLLM 0.26) and interior span (same content under a different wrapper; needs the scheduler re-consult patch in WorldFlowAI/vllm). See docs/QUICKSTART_VLLM.md.

In semantic_span_experimental the connector never advertises more than the donor KV it actually captured: spans trim to the stored donor window, a load that materializes zero layers fails loudly rather than decoding over uninitialized blocks, and every load leaves separate advertised and runtime_materialized audit records so reuse is only ever counted from the latter.

request_only_experimental defaults to exact-token-prefix materialization. The old zero-exact semantic proof behavior requires allow_non_identical_request_only=true or SEMBLEND_VLLM_ALLOW_NON_IDENTICAL_REQUEST_ONLY=1. Keep that flag limited to quality-gated validation experiments; it is not a production-safe substitute for a segmented/recompute engine path.

Safety Rules

The connector must not:

  • weaken exact prefix-cache semantics;
  • report semantic hits as computed tokens unless KV can actually be loaded;
  • publish non-identical semantic donor KV into vLLM's exact prefix cache;
  • treat non-identical semantic discovery as materializable unless an explicit validation flag is set and the run has separate quality gates;
  • cross model, tokenizer, adapter, or cache-salt namespaces;
  • fail inference because semantic lookup failed.

Repository Layout

src/semblend_vllm_connector/
  connector.py        vLLM KVConnectorBase_V1 implementation
  config.py           config/env parsing
  provider.py         provider protocol + local deterministic provider
  providers/
    semblend.py       lazy SemBlendPipeline adapter
  types.py            shared dataclasses/enums
  namespace.py        vLLM request namespace extraction

docs/
  ARCHITECTURE.md     detailed architecture and rollout plan
  SEMBLEND_PROVIDER.md
  VLLM_CONNECTOR_CONTRACT.md

examples/
  discovery_kv_transfer_config.json
  semblend_discovery_kv_transfer_config.json

Open Source Posture

This project follows the dynamic connector pattern used by mature vLLM KV cache projects: vLLM loads the connector from a Python module path, connector-specific settings live in kv_connector_extra_config, and unsafe materialization cases fail closed to normal vLLM prefill.

See:

Metadata

Release files for semblend-vllm-connector 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for semblend-vllm-connector 0.2.0
File Size Uploaded
semblend_vllm_connector-0.2.0.tar.gz 47.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for semblend-vllm-connector 0.2.0
File Interpreter ABI Platform
semblend_vllm_connector-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 75.8 kB

Release files / semblend_vllm_connector-0.2.0.tar.gz

Download URL semblend_vllm_connector-0.2.0.tar.gz
Size 47.2 kB
Tags Source
SHA-256 checksum
How to use checksums
5ad59a0450a17d0b2f12e18d1f4da0bd4944e0303cc22f38290423710c286b8e
BLAKE2b-256 checksum
How to use checksums
abb07266698f438e2f6b79dfea6436930f94512aa04f357c0694570dfcdc585e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release files / semblend_vllm_connector-0.2.0-py3-none-any.whl

Download URL semblend_vllm_connector-0.2.0-py3-none-any.whl
Size 28.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1a3302c0129bf232babac624f924504173f1f8aecdb17ef51caf82bf1593b148
BLAKE2b-256 checksum
How to use checksums
af92f0a2dd8cef65f1e6db7692cada93d4057f83f3eb45442c63dbdb83ab2f0e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release history Release notifications | RSS feed

0.2.8

2 release files

0.2.2

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page