Skip to main content

SemBlend vLLM Connector

CI License Python

vLLM KVConnector for SemBlend-backed semantic KV donor discovery.

This repo is the open-source adapter layer between vLLM and SemBlend.

SemBlend is a semantic KV reuse research library. It exists to evaluate when similar prompts may safely reuse or blend previously computed KV state. This connector exposes that work through vLLM's KVConnectorBase_V1 lifecycle.

Status

Experimental, with safe defaults. Verified paraphrase reuse in semantic_span_experimental mode has been validated end to end on stock vLLM 0.26 (Qwen2.5-7B, A10G) and through an llm-d gateway with semantic-affinity placement; start with docs/QUICKSTART_VLLM.md.

Default behavior is discovery-only:

  • exact vLLM prefix caching remains authoritative;
  • semantic lookup runs only after exact prefix coverage is insufficient;
  • the connector records donor hits, misses, and rejection reasons;
  • it returns (0, False) from get_num_new_matched_tokens() unless a configured materialization mode can prove a block-aligned exact token prefix or an explicitly opted-in isolated proof path;
  • normal vLLM execution continues on every provider error or unsupported case.

Install

From PyPI:

pip install "semblend-vllm-connector[semblend]"

Development:

pip install -e ".[semblend,dev]"

Run local checks:

make check

vLLM Configuration

Discovery-only mode:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-prefix-caching \
  --kv-transfer-config '{
    "kv_connector": "SemBlendVllmConnector",
    "kv_connector_module_path": "semblend_vllm_connector.connector",
    "kv_role": "kv_both",
    "kv_load_failure_policy": "recompute",
    "kv_connector_extra_config": {
      "mode": "discovery_only",
      "provider": "local",
      "min_prompt_tokens": 512,
      "min_similarity": 0.70
    }
  }'

SemBlend provider mode:

{
  "kv_connector": "SemBlendVllmConnector",
  "kv_connector_module_path": "semblend_vllm_connector.connector",
  "kv_role": "kv_both",
  "kv_load_failure_policy": "recompute",
  "kv_connector_extra_config": {
    "mode": "discovery_only",
    "provider": "semblend",
    "min_prompt_tokens": 512,
    "min_similarity": 0.70,
    "min_reuse_ratio": 0.50,
    "embedder_type": "minilm",
    "model_id": "meta-llama/Llama-3.1-8B-Instruct"
  }
}

Equivalent JSON examples live in examples/.

Optional audit stream for reproducible validation:

{
  "kv_connector_extra_config": {
    "audit_path": "/tmp/semblend-vllm-audit.jsonl",
    "log_decisions": true
  }
}

The audit file is JSONL. Runtime KV reuse should be counted only from runtime_materialized events. Semantic lookup hits and advertised loads are reported separately so benchmark runners can distinguish discovery from backend-confirmed materialization.

Modes

Mode Positive matched tokens? Purpose
discovery_only No Safe telemetry and workload qualification.
exact_prefix Only with engine-valid exact block refs Future safe materialization path.
request_only_experimental Yes, exact-token-prefix blocks by default Isolated validation mode; run with vLLM prefix caching disabled.
segmented_experimental Not enabled in this repo yet Requires segmented/sparse execution and recompute-boundary support.
semantic_span_experimental Yes, block-aligned donor spans with RoPE re-rotation Semantic reuse of non-identical prompts. Two lanes: verified paraphrase (whole-span serve gated by a fail-closed fact check; works on stock vLLM 0.26) and interior span (same content under a different wrapper; needs the scheduler re-consult patch in WorldFlowAI/vllm). See docs/QUICKSTART_VLLM.md.

In semantic_span_experimental the connector never advertises more than the donor KV it actually captured: spans trim to the stored donor window, a load that materializes zero layers fails loudly rather than decoding over uninitialized blocks, and every load leaves separate advertised and runtime_materialized audit records so reuse is only ever counted from the latter.

request_only_experimental defaults to exact-token-prefix materialization. The old zero-exact semantic proof behavior requires allow_non_identical_request_only=true or SEMBLEND_VLLM_ALLOW_NON_IDENTICAL_REQUEST_ONLY=1. Keep that flag limited to quality-gated validation experiments; it is not a production-safe substitute for a segmented/recompute engine path.

Safety Rules

The connector must not:

  • weaken exact prefix-cache semantics;
  • report semantic hits as computed tokens unless KV can actually be loaded;
  • publish non-identical semantic donor KV into vLLM's exact prefix cache;
  • treat non-identical semantic discovery as materializable unless an explicit validation flag is set and the run has separate quality gates;
  • cross model, tokenizer, adapter, or cache-salt namespaces;
  • fail inference because semantic lookup failed.

The blocks a semantic load fills are therefore evicted from vLLM's exact prefix cache. evict_filled_blocks_from_prefix_cache: false switches that off and is the one supported way to break the third rule above: it exists so a measurement can quantify what the eviction removes, and it lets approximate KV be served to a later exact match. Leave it at its default (true) anywhere else. Compare the two runs on prefix_cache_distinct_blocks_evicted against prefix_cache_distinct_blocks_left_cached; the per-pass counts beside them are not comparable across the arms (audit contract).

Repository Layout

src/semblend_vllm_connector/
  connector.py        vLLM KVConnectorBase_V1 implementation
  config.py           config/env parsing
  provider.py         provider protocol + local deterministic provider
  providers/
    semblend.py       lazy SemBlendPipeline adapter
  types.py            shared dataclasses/enums
  namespace.py        vLLM request namespace extraction

docs/
  ARCHITECTURE.md     detailed architecture and rollout plan
  SEMBLEND_PROVIDER.md
  VLLM_CONNECTOR_CONTRACT.md

examples/
  discovery_kv_transfer_config.json
  semblend_discovery_kv_transfer_config.json

Open Source Posture

This project follows the dynamic connector pattern used by mature vLLM KV cache projects: vLLM loads the connector from a Python module path, connector-specific settings live in kv_connector_extra_config, and unsafe materialization cases fail closed to normal vLLM prefill.

See:

Metadata

Release files for semblend-vllm-connector 0.2.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for semblend-vllm-connector 0.2.8
File Size Uploaded
semblend_vllm_connector-0.2.8.tar.gz 207.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for semblend-vllm-connector 0.2.8
File Interpreter ABI Platform
semblend_vllm_connector-0.2.8-py3-none-any.whl Python 3 none any Details

Total release size: 297.0 kB

Release files / semblend_vllm_connector-0.2.8.tar.gz

Download URL semblend_vllm_connector-0.2.8.tar.gz
Size 207.8 kB
Tags Source
SHA-256 checksum
How to use checksums
6683fdff79d88203c29637f173914a6a35a8c428d0823f1aa481476f3187f444
BLAKE2b-256 checksum
How to use checksums
9a27950d6a4706bee8cd267dd1d5fe138b6d73f41e8780131d34c0300ba46d86
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release files / semblend_vllm_connector-0.2.8-py3-none-any.whl

Download URL semblend_vllm_connector-0.2.8-py3-none-any.whl
Size 89.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a3a7eb8e138a5f8a1b5faefe633427ed6f6f9a11f05fc75d5f29a1f3293ed588
BLAKE2b-256 checksum
How to use checksums
35c72d34a09add2b851b507905ce78d96c42db9fe98d0911e80e143c919caef4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release history Release notifications | RSS feed

This release

0.2.8 This release

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page