tetherto-qvac-sdk — Python SDK
The Python client for QVAC: local-first, P2P AI inference (LLM completion,
embeddings, transcription, TTS, OCR, translation, diffusion, audio generation,
VLA, …) through the same worker and the same contract as the TypeScript
@qvac/sdk. Asyncio-native.
The package version tracks the @qvac/sdk version it speaks (e.g. 0.15.0), so
a given python-sdk release makes clear exactly which SDK it targets.
Install
The fastest path is a self-contained install — one command, no separate
worker, no Node.js. It pulls a per-platform wheel that bundles the QVAC worker
(the @qvac/sdk runtime) and the Bare runtime, from the matching GitHub
release (these wheels are far too large for PyPI, so they live as release
assets):
# Replace <version> with the release you want, e.g. sdk-v0.17.0:
pip install tetherto-qvac-sdk \
-f https://github.com/tetherto/qvac/releases/expanded_assets/sdk-v<version>
pip selects the bundled wheel for your platform and pulls the pure-Python deps
(pydantic, bare-rpc, compact-encoding) from PyPI. Then
async with Client() as client: ... just works — the bundled worker is found
automatically, zero configuration. Optional extras: vla (numpy),
notebook (numpy + pandas).
Bundled wheels ship for darwin-arm64, linux-x64, linux-arm64, win32-x64. On any other platform pip falls back to the thin wheel below.
Thin install (any platform)
If there's no bundled wheel for your platform, or you'd rather run a shared system worker, install the thin client from PyPI and provide the worker separately:
pip install tetherto-qvac-sdk
Then install the worker once (either route needs Node.js):
# via Python — fetches the exact worker version this package speaks and caches
# it under ~/.cache/qvac/worker/<version>:
python -m tetherto.qvac_sdk install-worker
# or via npm — install the @qvac/sdk version that MATCHES this package (they
# share a version; Client() warns on a mismatch):
npm install -g @qvac/sdk@<your tetherto-qvac-sdk version>
Client() then finds the worker automatically (see "Worker resolution" below
for the full lookup order, including pointing QVAC_SDK_DIR at a locally-built
@qvac/sdk for development). The Python route pins the worker to the exact
version this package was generated against, so you never track it yourself.
Upgrading
To upgrade and keep the bundled worker, re-run the install with -f
pointing at the new release tag:
pip install -U tetherto-qvac-sdk \
-f https://github.com/tetherto/qvac/releases/expanded_assets/sdk-v<newversion>
A plain pip install -U tetherto-qvac-sdk (no -f, or a stale tag) upgrades to
the thin PyPI wheel: pip applies no priority between locations and just
takes the highest version it can see, which on PyPI is the py3-none-any
wheel — you'd then need install-worker. To switch an already-installed thin
build to the bundled wheel at the same version, add --force-reinstall
(pip treats a version as satisfied regardless of which wheel is installed).
Quickstart
import asyncio
from tetherto.qvac_sdk import Client, load_model, completion, unload_model
from tetherto.qvac_sdk.models import LLAMA_3_2_1B_INST_Q4_0
async def main():
async with Client() as client:
t = client.transport
model_id = await load_model(t, model_src=LLAMA_3_2_1B_INST_Q4_0)
run = completion(
t,
model_id=model_id,
history=[{"role": "user", "content": "Explain quantum computing in one sentence"}],
)
async for event in run.events:
if event.type == "contentDelta":
print(event.text, end="", flush=True)
print()
await unload_model(t, model_id)
asyncio.run(main())
More in examples/ — one runnable example per major capability
(completion events / tools / worker-orchestrated tools, cancel, embeddings,
translation, transcription, TTS, OCR, audio generation, registry queries, model
info, logging, VLA, plugins), mirroring packages/sdk/examples.
The public API (tetherto.qvac_sdk)
Everything you normally need is re-exported flat from tetherto.qvac_sdk; model constants
live in tetherto.qvac_sdk.models.
Client— starts/owns a worker connection;client.transportis passed to every call.- Ergonomic wrappers (kwargs + typed results):
load_model,unload_model,completion,translate,cancel,delete_cache,invoke_plugin/invoke_plugin_stream,model_registry_list/_search/_get_model. (Tool calling iscompletion(tools=...); the worker-orchestrated loop is an advanced path attetherto.qvac_sdk._completion.completion_orchestrate, not flat-public.) - Result types:
CompletionRun(.events,.final),CompletionFinal,ToolCall,TranslateRun. - Generated method stubs for every other contract method (
embed,transcribe,text_to_speech,ocr_stream,audio_gen_stream,diffusion_stream,classify,get_model_info,download_asset, …), each taking a typed request model. - Request/response models + enums (
LoadModelRequest,ModelType, …), also available in full fromtetherto.qvac_sdk.schemas. - Errors (
QvacError,RPCError,InferenceCancelledError,ContextOverflowError, …) forexcept/isinstance. - Logging:
logging_stream,subscribe_server_logs,SDK_LOG_ID,SDK_ALL_LOG_ID. VLA:vla,vla_hparams,vla_preprocess_image,vla_pad_state. - Notebook facade:
tetherto.qvac_sdk.notebook.SyncClient— synchronous, numpy/pandas returns, live in-cell streaming.
The raw generated method stubs are also in tetherto.qvac_sdk.methods, and the pydantic
models in tetherto.qvac_sdk.schemas, if you prefer the explicit modules.
Audio generation currently uses the raw audio_gen_stream stub rather than an
ergonomic Python audio_gen() wrapper. Build an AudioGenStreamRequest, iterate
the progress and base64 PCM frames, and assemble the output audio. See
examples/audiogen.py.
Intentional divergences from @qvac/sdk
Two top-level JS/TS facades are deliberately not ported; Python uses the ecosystem equivalent instead.
getLogger— use the standard libraryloggingmodule. Alogging.Handleris theLogTransportequivalent, and the log-stream surface that has no stdlib counterpart (logging_stream,subscribe_server_logs,SDK_LOG_ID,SDK_ALL_LOG_ID) is already exported above.profiler(process-wide aggregation +exportTable/exportJSON) — use the per-callprofiled_calland the__profilingenvelope helpers intetherto.qvac_sdk.profiling, aggregating the returnedProfilingReports yourself. The global profiler is not a Client-API capability and would only collect useful data after instrumenting the full streaming client.
Notebook / data science
For notebooks and REPLs, tetherto.qvac_sdk.notebook.SyncClient runs the async
client on a background thread so every call is plain and blocking (no await),
with numpy/pandas returns and live in-cell streaming. Needs the notebook
extra (pip install "tetherto-qvac-sdk[notebook]"):
from tetherto.qvac_sdk.notebook import SyncClient
from tetherto.qvac_sdk.models import EMBEDDINGGEMMA_300M_Q4_0
with SyncClient() as client:
m = client.load_model(model_src=EMBEDDINGGEMMA_300M_Q4_0)
vec = client.embed(m, "hello") # 1-D numpy array
df = client.embed_frame(m, ["a", "b", "c"]) # pandas DataFrame, indexed by text
embed returns numpy arrays, embed_frame a DataFrame, completion streams
live and returns the text, and transcribe/text_to_speech round-trip audio
as numpy. Full examples: examples/notebook.ipynb
(Jupyter) and examples/notebook.py (script).
Configuring the SDK
Client(config=...) (and BareRpcTransport(config=...)) sends an SDK
config to the worker on connect — the same QvacConfig the TypeScript client
applies (cacheDirectory, loggerLevel, swarmRelays, per-device plugin
defaults, …):
async with Client(config={"cacheDirectory": "/data/qvac-models", "loggerLevel": "warn"}) as client:
...
cacheDirectory (must be absolute) relocates where models are stored — handy
for a shared/mounted cache. As a shortcut, the QVAC_CACHE_DIR env var sets
cacheDirectory without code (used by CI to point at a warm model cache).
Worker resolution
Client() locates a worker by trying, in order: explicit worker_path /
bare_path (or QVAC_WORKER_PATH / QVAC_BARE_PATH); a self-contained
bundled wheel (tetherto/qvac_sdk/_bundle/, zero config); sdk_dir / QVAC_SDK_DIR
pointing at an installed @qvac/sdk; a worker fetched by python -m tetherto.qvac_sdk install-worker (into ~/.cache/qvac/worker/<version>, overridable with
QVAC_WORKER_HOME); then a global npm install -g @qvac/sdk. The fetched
worker is version-locked to this package.
Staying in sync with @qvac/sdk (drift avoidance)
The rule: generate what can be generated; put un-generatable behaviour in the worker; hand-write per language only what is irreducibly client-side, and never trust it to match — guard it.
- Generated from the one contract. The pydantic models + typed method stubs,
the model-type resolution maps (
_generated/model_type_maps.py), the error-code registries (_generated/error_codes.py), and the pinned SDK version (_generated/sdk_version.py) are all generated from../sdk/contract/**.generate.py --check(and the SDK'scontract:check) fail CI if either side drifts. - Behaviour lives in the worker.
translatesource-language detection and the completion tool-loop (completionOrchestrate) run in the worker, so Python and JS share one implementation instead of two that diverge (they did: Python once usedlingua, JS@qvac/langdetect-text). - Version lock-step. This package's version is the
@qvac/sdkversion it was generated against. - Conformance corpus.
../sdk/e2e/conformance/cases.jsonis run by both a JS runner andtests/test_conformance.py, so the two clients are diffed against the same cases.
The irreducibly-client-side code (typed error classes, numpy marshaling, stream assembly, the notebook facade) is the only hand-written surface mirroring the JS client, and it's covered by the conformance corpus + real-worker tests rather than trusted to match.
Development
python3 -m venv .venv
.venv/bin/pip install -e ".[gen,dev]"
.venv/bin/python3 scripts/generate.py # regenerate from ../sdk/contract
.venv/bin/python3 -m pytest # unit + (with a built worker) real-model e2e
Format, lint, typecheck, and the generation check (all run in CI):
.venv/bin/python3 scripts/generate.py --check
.venv/bin/python3 -m black --check src/tetherto/qvac_sdk scripts/ tests/
.venv/bin/python3 -m ruff check src/tetherto/qvac_sdk scripts/ tests/
.venv/bin/python3 -m mypy -p tetherto.qvac_sdk && .venv/bin/python3 -m mypy scripts tests
Real-model tests spawn a worker (packages/sdk built via bun run build, or
QVAC_POC_SDK_DIR) and otherwise skip. generate.py runs black + ruff --fix --select I (with the package config) on its own output, so a fresh
regeneration already passes the checks above.
Release files for tetherto-qvac-sdk 0.18.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tetherto_qvac_sdk-0.18.0.tar.gz | 218.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tetherto_qvac_sdk-0.18.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 379.0 kB
Release files / tetherto_qvac_sdk-0.18.0.tar.gz
| Download URL | tetherto_qvac_sdk-0.18.0.tar.gz |
|---|---|
| Size | 218.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cb0cb853f39a2f1d4b20478d8071371e317b69ebb67375e02c61f000af353a9c
|
|
BLAKE2b-256 checksum How to use checksums |
b5b2ed02f0f9cdabd72f17cfa43ae6cf3a671535b03fee6e9b95fdf1aefd7d40
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency logRelease files / tetherto_qvac_sdk-0.18.0-py3-none-any.whl
| Download URL | tetherto_qvac_sdk-0.18.0-py3-none-any.whl |
|---|---|
| Size | 160.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c1419962f0c6b49801b93ca0f2b185acbc5e4fce35648a96d57e3787a275bbd9
|
|
BLAKE2b-256 checksum How to use checksums |
a24e72f7cab67e37ab072ff06f5dcf69020a7fde290226105f661ddb7cd0688b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency log