SemBlend vLLM Connector
vLLM KVConnector for SemBlend-backed semantic KV donor discovery.
This repo is the open-source adapter layer between vLLM and SemBlend.
SemBlend is a semantic KV reuse research library. It exists to evaluate when
similar prompts may safely reuse or blend previously computed KV state. This
connector exposes that work through vLLM's KVConnectorBase_V1
lifecycle.
Status
Experimental, with safe defaults. Verified paraphrase reuse in
semantic_span_experimental mode has been validated end to end on stock vLLM
0.26 (Qwen2.5-7B, A10G) and through an llm-d gateway with semantic-affinity
placement; start with docs/QUICKSTART_VLLM.md.
Default behavior is discovery-only:
- exact vLLM prefix caching remains authoritative;
- semantic lookup runs only after exact prefix coverage is insufficient;
- the connector records donor hits, misses, and rejection reasons;
- it returns
(0, False)fromget_num_new_matched_tokens()unless a configured materialization mode can prove a block-aligned exact token prefix or an explicitly opted-in isolated proof path; - normal vLLM execution continues on every provider error or unsupported case.
Install
From PyPI:
pip install "semblend-vllm-connector[semblend]"
Development:
pip install -e ".[semblend,dev]"
Run local checks:
make check
vLLM Configuration
Discovery-only mode:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--enable-prefix-caching \
--kv-transfer-config '{
"kv_connector": "SemBlendVllmConnector",
"kv_connector_module_path": "semblend_vllm_connector.connector",
"kv_role": "kv_both",
"kv_load_failure_policy": "recompute",
"kv_connector_extra_config": {
"mode": "discovery_only",
"provider": "local",
"min_prompt_tokens": 512,
"min_similarity": 0.70
}
}'
SemBlend provider mode:
{
"kv_connector": "SemBlendVllmConnector",
"kv_connector_module_path": "semblend_vllm_connector.connector",
"kv_role": "kv_both",
"kv_load_failure_policy": "recompute",
"kv_connector_extra_config": {
"mode": "discovery_only",
"provider": "semblend",
"min_prompt_tokens": 512,
"min_similarity": 0.70,
"min_reuse_ratio": 0.50,
"embedder_type": "minilm",
"model_id": "meta-llama/Llama-3.1-8B-Instruct"
}
}
Equivalent JSON examples live in examples/.
Optional audit stream for reproducible validation:
{
"kv_connector_extra_config": {
"audit_path": "/tmp/semblend-vllm-audit.jsonl",
"log_decisions": true
}
}
The audit file is JSONL. Runtime KV reuse should be counted only from
runtime_materialized events. Semantic lookup hits and advertised loads are
reported separately so benchmark runners can distinguish discovery from
backend-confirmed materialization.
Modes
| Mode | Positive matched tokens? | Purpose |
|---|---|---|
discovery_only |
No | Safe telemetry and workload qualification. |
exact_prefix |
Only with engine-valid exact block refs | Future safe materialization path. |
request_only_experimental |
Yes, exact-token-prefix blocks by default | Isolated validation mode; run with vLLM prefix caching disabled. |
segmented_experimental |
Not enabled in this repo yet | Requires segmented/sparse execution and recompute-boundary support. |
semantic_span_experimental |
Yes, block-aligned donor spans with RoPE re-rotation | Semantic reuse of non-identical prompts. Two lanes: verified paraphrase (whole-span serve gated by a fail-closed fact check; works on stock vLLM 0.26) and interior span (same content under a different wrapper; needs the scheduler re-consult patch in WorldFlowAI/vllm). See docs/QUICKSTART_VLLM.md. |
In semantic_span_experimental the connector never advertises more than the
donor KV it actually captured: spans trim to the stored donor window, a load
that materializes zero layers fails loudly rather than decoding over
uninitialized blocks, and every load leaves separate advertised and
runtime_materialized audit records so reuse is only ever counted from the
latter.
request_only_experimental defaults to exact-token-prefix materialization. The
old zero-exact semantic proof behavior requires
allow_non_identical_request_only=true or
SEMBLEND_VLLM_ALLOW_NON_IDENTICAL_REQUEST_ONLY=1. Keep that flag limited to
quality-gated validation experiments; it is not a production-safe substitute for a
segmented/recompute engine path.
Safety Rules
The connector must not:
- weaken exact prefix-cache semantics;
- report semantic hits as computed tokens unless KV can actually be loaded;
- publish non-identical semantic donor KV into vLLM's exact prefix cache;
- treat non-identical semantic discovery as materializable unless an explicit validation flag is set and the run has separate quality gates;
- cross model, tokenizer, adapter, or cache-salt namespaces;
- fail inference because semantic lookup failed.
The blocks a semantic load fills are therefore evicted from vLLM's exact prefix
cache. evict_filled_blocks_from_prefix_cache: false switches that off and is
the one supported way to break the third rule above: it exists so a measurement
can quantify what the eviction removes, and it lets approximate KV be served to
a later exact match. Leave it at its default (true) anywhere else. Compare
the two runs on prefix_cache_distinct_blocks_evicted against
prefix_cache_distinct_blocks_left_cached; the per-pass counts beside them are
not comparable across the arms
(audit contract).
Repository Layout
src/semblend_vllm_connector/
connector.py vLLM KVConnectorBase_V1 implementation
config.py config/env parsing
provider.py provider protocol + local deterministic provider
providers/
semblend.py lazy SemBlendPipeline adapter
types.py shared dataclasses/enums
namespace.py vLLM request namespace extraction
docs/
ARCHITECTURE.md detailed architecture and rollout plan
SEMBLEND_PROVIDER.md
VLLM_CONNECTOR_CONTRACT.md
examples/
discovery_kv_transfer_config.json
semblend_discovery_kv_transfer_config.json
Open Source Posture
This project follows the dynamic connector pattern used by mature vLLM KV cache
projects: vLLM loads the connector from a Python module path, connector-specific
settings live in kv_connector_extra_config, and unsafe materialization cases
fail closed to normal vLLM prefill.
See:
Metadata
Release files for semblend-vllm-connector 0.2.8
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| semblend_vllm_connector-0.2.8.tar.gz | 207.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| semblend_vllm_connector-0.2.8-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 297.0 kB
Release files / semblend_vllm_connector-0.2.8.tar.gz
| Download URL | semblend_vllm_connector-0.2.8.tar.gz |
|---|---|
| Size | 207.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6683fdff79d88203c29637f173914a6a35a8c428d0823f1aa481476f3187f444
|
|
BLAKE2b-256 checksum How to use checksums |
9a27950d6a4706bee8cd267dd1d5fe138b6d73f41e8780131d34c0300ba46d86
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|
Release files / semblend_vllm_connector-0.2.8-py3-none-any.whl
| Download URL | semblend_vllm_connector-0.2.8-py3-none-any.whl |
|---|---|
| Size | 89.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a3a7eb8e138a5f8a1b5faefe633427ed6f6f9a11f05fc75d5f29a1f3293ed588
|
|
BLAKE2b-256 checksum How to use checksums |
35c72d34a09add2b851b507905ce78d96c42db9fe98d0911e80e143c919caef4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|