SemBlend vLLM Connector
vLLM KVConnector for SemBlend-backed semantic KV donor discovery.
This repo is the open-source adapter layer between vLLM and SemBlend.
SemBlend is a semantic KV reuse research library. It exists to evaluate when
similar prompts may safely reuse or blend previously computed KV state. This
connector exposes that work through vLLM's KVConnectorBase_V1
lifecycle.
Status
Experimental, with safe defaults. Verified paraphrase reuse in
semantic_span_experimental mode has been validated end to end on stock vLLM
0.26 (Qwen2.5-7B, A10G) and through an llm-d gateway with semantic-affinity
placement; start with docs/QUICKSTART_VLLM.md.
Default behavior is discovery-only:
- exact vLLM prefix caching remains authoritative;
- semantic lookup runs only after exact prefix coverage is insufficient;
- the connector records donor hits, misses, and rejection reasons;
- it returns
(0, False)fromget_num_new_matched_tokens()unless a configured materialization mode can prove a block-aligned exact token prefix or an explicitly opted-in isolated proof path; - normal vLLM execution continues on every provider error or unsupported case.
Install
From PyPI:
pip install "semblend-vllm-connector[semblend]"
Development:
pip install -e ".[semblend,dev]"
Run local checks:
make check
vLLM Configuration
Discovery-only mode:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--enable-prefix-caching \
--kv-transfer-config '{
"kv_connector": "SemBlendVllmConnector",
"kv_connector_module_path": "semblend_vllm_connector.connector",
"kv_role": "kv_both",
"kv_load_failure_policy": "recompute",
"kv_connector_extra_config": {
"mode": "discovery_only",
"provider": "local",
"min_prompt_tokens": 256,
"min_similarity": 0.70
}
}'
SemBlend provider mode:
{
"kv_connector": "SemBlendVllmConnector",
"kv_connector_module_path": "semblend_vllm_connector.connector",
"kv_role": "kv_both",
"kv_load_failure_policy": "recompute",
"kv_connector_extra_config": {
"mode": "discovery_only",
"provider": "semblend",
"min_prompt_tokens": 256,
"min_similarity": 0.70,
"min_reuse_ratio": 0.50,
"embedder_type": "minilm",
"model_id": "meta-llama/Llama-3.1-8B-Instruct"
}
}
Equivalent JSON examples live in examples/.
Optional audit stream for reproducible validation:
{
"kv_connector_extra_config": {
"audit_path": "/tmp/semblend-vllm-audit.jsonl",
"log_decisions": true
}
}
The audit file is JSONL. Runtime KV reuse should be counted only from
runtime_materialized events. Semantic lookup hits and advertised loads are
reported separately so benchmark runners can distinguish discovery from
backend-confirmed materialization.
Modes
| Mode | Positive matched tokens? | Purpose |
|---|---|---|
discovery_only |
No | Safe telemetry and workload qualification. |
exact_prefix |
Only with engine-valid exact block refs | Future safe materialization path. |
request_only_experimental |
Yes, exact-token-prefix blocks by default | Isolated validation mode; run with vLLM prefix caching disabled. |
segmented_experimental |
Not enabled in this repo yet | Requires segmented/sparse execution and recompute-boundary support. |
semantic_span_experimental |
Yes, block-aligned donor spans with RoPE re-rotation | Semantic reuse of non-identical prompts. Two lanes: verified paraphrase (whole-span serve gated by a fail-closed fact check; works on stock vLLM 0.26) and interior span (same content under a different wrapper; needs the scheduler re-consult patch in WorldFlowAI/vllm). See docs/QUICKSTART_VLLM.md. |
In semantic_span_experimental the connector never advertises more than the
donor KV it actually captured: spans trim to the stored donor window, a load
that materializes zero layers fails loudly rather than decoding over
uninitialized blocks, and every load leaves separate advertised and
runtime_materialized audit records so reuse is only ever counted from the
latter.
request_only_experimental defaults to exact-token-prefix materialization. The
old zero-exact semantic proof behavior requires
allow_non_identical_request_only=true or
SEMBLEND_VLLM_ALLOW_NON_IDENTICAL_REQUEST_ONLY=1. Keep that flag limited to
quality-gated validation experiments; it is not a production-safe substitute for a
segmented/recompute engine path.
Safety Rules
The connector must not:
- weaken exact prefix-cache semantics;
- report semantic hits as computed tokens unless KV can actually be loaded;
- publish non-identical semantic donor KV into vLLM's exact prefix cache;
- treat non-identical semantic discovery as materializable unless an explicit validation flag is set and the run has separate quality gates;
- cross model, tokenizer, adapter, or cache-salt namespaces;
- fail inference because semantic lookup failed.
Repository Layout
src/semblend_vllm_connector/
connector.py vLLM KVConnectorBase_V1 implementation
config.py config/env parsing
provider.py provider protocol + local deterministic provider
providers/
semblend.py lazy SemBlendPipeline adapter
types.py shared dataclasses/enums
namespace.py vLLM request namespace extraction
docs/
ARCHITECTURE.md detailed architecture and rollout plan
SEMBLEND_PROVIDER.md
VLLM_CONNECTOR_CONTRACT.md
examples/
discovery_kv_transfer_config.json
semblend_discovery_kv_transfer_config.json
Open Source Posture
This project follows the dynamic connector pattern used by mature vLLM KV cache
projects: vLLM loads the connector from a Python module path, connector-specific
settings live in kv_connector_extra_config, and unsafe materialization cases
fail closed to normal vLLM prefill.
See:
Metadata
Release files for semblend-vllm-connector 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| semblend_vllm_connector-0.2.1.tar.gz | 50.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| semblend_vllm_connector-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 79.9 kB
Release files / semblend_vllm_connector-0.2.1.tar.gz
| Download URL | semblend_vllm_connector-0.2.1.tar.gz |
|---|---|
| Size | 50.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0076a8d58bf3dc97254942e8c29643bb9f3b5feb90a662325452e1d37c2e03b0
|
|
BLAKE2b-256 checksum How to use checksums |
b027697e0cf22b2ebb9b8c9e751bba0710692106d4a9c3dfd92a9b5063ccd4be
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|
Release files / semblend_vllm_connector-0.2.1-py3-none-any.whl
| Download URL | semblend_vllm_connector-0.2.1-py3-none-any.whl |
|---|---|
| Size | 29.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6324fffcf80c25e6102141d4a5eb749ded2f25782312c778be4e92b91a430b73
|
|
BLAKE2b-256 checksum How to use checksums |
70d1015a10d5ad4f67f3c24ba3370cf184fed612b10207a891cf0ed439c2d8cc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|