QEV
Your evidence. Your criteria. A decision with probabilities.
Model on Hugging Face · Dataset
QEV is an open multimodal decision model inspired by LAYA and built with the text and vision backbone of Qwen3.5-2B. Give it text, a photo, or both, then describe the decision you need. It returns candidate probabilities, a typed answer and an abstention signal in one batched backbone forward, with zero generated answer tokens.
LAYA's request-defined typed decisions motivated the interface. QEV adds visual evidence through Qwen's existing multimodal backbone, learned language adapters, option readouts and a condition-modulated binding head. The Qwen vision encoder is frozen. LAYA weights are not embedded in this checkpoint, and the current training recipe is supervised adaptation with calibration. RLCD remains a research direction. QEV is independently maintained; upstream model attribution is preserved.
Install and run
The QEV 0.2.0 Python SDK includes the English playground, six licensed example photographs, image resolution controls, and the inference API. Use Python 3.12.
pip
python -m pip install https://huggingface.co/ken-jo/qev/resolve/main/runtime/qev-0.2.0-py3-none-any.whl
qev playground
Open http://127.0.0.1:7860. The first launch downloads the pinned QEV adaptation and
Qwen3.5-2B backbone (about 4.6 GB), then loads the model. Later launches reuse the cache.
CUDA is used when available; --device cpu and --device cuda select a device explicitly.
For availability of the shorter pip install qev command, see
PyPI publication status.
uv
Run the published wheel without cloning the repository:
uvx --python 3.12 --from https://huggingface.co/ken-jo/qev/resolve/main/runtime/qev-0.2.0-py3-none-any.whl qev playground
Or use the source checkout and its frozen CUDA dependency lock:
git clone https://github.com/ken-jo/qev.git
cd qev
uv sync --frozen
uv run qev playground
For NVIDIA CUDA 12.8 wheels with uvx, add
--torch-backend cu128 before --from. With pip, install the
matching CUDA-enabled PyTorch and torchvision first if the default installation is CPU-only.
Python SDK
from qev import DecisionRequest, load
model = load() # Downloads on first use; reuses the model cache afterwards.
request = DecisionRequest.model_validate({
"state": {"text": "I was charged twice. Please refund the duplicate payment."},
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {"billing": "Payments and refunds", "technical": "Software faults"},
}
},
})
print(model.predict(request)["answers"])
For JSON files: qev predict --request examples/request.json. For an HTTP API:
qev serve --image-root examples. Both prepare the model on first use. qev download
can fetch weights in advance; --offline requires a complete cache. Set QEV_HOME for
QEV's persistent data directory or QEV_CACHE_DIR for the Hugging Face cache.
The Python distribution contains code, UI and example photos. Model tensors are downloaded
separately. The inference weights and 38 internal veyra modules remain the evaluated
QEV 0.1.1 model; SDK 0.2.0 adds installation and application features.
See the playground guide
and API schema.
Decisions you define at request time
Route a support message, assess an ordered severity level, or ask whether a photo meets a written condition. Change the candidate descriptions to change the task. The model scores the choices supplied with the request; its output slots do not represent a fixed catalog of classes.
These are intended uses for domain evaluation. Published results below establish the current scope, including failures in visual reasoning and unfamiliar tasks.
What you can ask
| Type | Request-specific definition | Output |
|---|---|---|
choice |
2-16 candidate descriptions | Candidate probabilities and selected label |
score |
2-16 ordered level descriptions | Level probabilities and expected level |
noul |
A proposition, with true/false criteria | True/false probabilities |
Accepts text, one image, or both; 1-4 questions per request. Candidate meanings are supplied at inference time. This does not imply reliable generalization to every unseen task.
flowchart LR
E[Text and optional image] --> Q[Qwen3.5-2B with language LoRA]
C[Question and candidate descriptions] --> Q
Q --> R[Option readout and condition binding]
R --> P[Type-specific calibration]
P --> A[Probabilities, typed answer, abstention]
QEV and LAYA on the same inputs
| English test | LAYA English | LAYA Typed Decisions | QEV 0.1.1 |
|---|---|---|---|
| Typed decisions, 2,000 questions | 36.05% | 76.95% | 77.00% |
| News topic, 400 examples | 95.00% | 95.25% | 82.50% |
| Emotion, 400 examples | 58.75% | 60.00% | 50.25% |
Typed-decisions is an adapted benchmark for QEV and the specialist. Their one-question accuracy difference does not establish an advantage (paired 95% interval: -1.85 to +1.95 percentage points). News and emotion are zero-shot relative to QEV's audited adaptation data; unknown backbone pretraining overlap remains possible. LAYA reports news in its training mix and emotion held out. All models received the same inputs, without truncation.
LAYA is smaller and faster in this comparison. On typed-decisions, its specialist has lower probability errors and 23.71 ms resident p50 versus QEV's 77.70 ms on the same GPU. QEV adds image input; that capability has its own evaluations and limitations. See the full protocol, probability metrics and results.
A photo, three typed answers
The release includes a reproducible photograph example: one plastic bottle, six material candidates, a three-level handling policy, and a true/false proposition.
| Type | Recorded output |
|---|---|
choice |
Plastic: 0.9361 probability |
score |
Expected policy level: 0.4254 on a 0-2 scale |
noul |
Glass, paper or plastic: 0.7903 probability true |
One batched forward, zero generated answer tokens. These are actual outputs on a previously inspected verification fixture, not an independent accuracy estimate. Photo: TrashNet, Gary Thung, MIT. Request, output, image and reproduction.
Image and workflow evaluation
| Evaluation | Result | What it covers |
|---|---|---|
| Fresh procedural workflows | 69.38% | 1,440 questions, 480 groups, 3 authored families |
| Fresh CIFAR-10 photograph guard | 95.83% | 600 low-resolution image questions |
| Fresh SNLI | 87.78% | 450 questions |
| Fresh BANKING77 | 85.50% | 462 sampled eight-candidate questions, not standard 77-way classification |
| Official typed-decisions regression | 77.00% | 2,000 questions; 30.80% coverage under abstention |
| Resident local photo HTTP p95 | 114.94 ms | RTX 4060 Ti 8 GB; serial, one photo/question, six candidates |
The timing excludes loading, WAN transport and concurrency. It is not a general latency SLA. On the shifted final population, accepted uncertain requests had 45.57% expected error; overall ECE was 12.58%. Calibration does not guarantee correctness.
Visual reasoning remains limited: 0 of 72 exploratory 2048 games reached 2048. A narrow balanced board-cell diagnostic scored 5/28 for images and 21/28 for text. These failures are published alongside the successful measurements. Read the model card and evaluation report before using it.
What you download
| Artifact | Hosted on | Purpose |
|---|---|---|
| QEV adaptation and decision heads | Hugging Face ken-jo/qev |
The learned QEV weights and calibration |
| Qwen3.5-2B backbone | Hugging Face Qwen/Qwen3.5-2B |
The pinned upstream text and vision model |
qev Python SDK |
Release wheel; see publication status above | Loading, typed inference, downloads, serving and playground |
| Research and application source | GitHub ken-jo/qev |
Training scripts, evaluations and local interfaces |
Installing the SDK installs code, the playground, samples and dependencies. First use
fetches the two weight components; qev download can prepare them in advance. The SDK's package size is not the model's size.
Size and precision
The inference backbone has 2.213B parameters. The checkpoint adds 7.992M stored adapter/readout parameters in a 32.01 MB safetensors file. GPU inference uses a BF16 backbone and FP32 decision readouts; the CPU path uses FP32. LoRA is merged into the backbone for inference. Upstream weights download separately (about 4.55 GB). See exact counts and tensor types.
Playground and API
qev playground starts the packaged English interface. It includes text/image presets,
six sample photographs, choice/score/noul, image-resolution choices, candidate probabilities,
abstention and the exact request/response JSON. It runs locally; no public Space is created.
qev playground --host 0.0.0.0 --port 7860
qev serve --image-root examples --host 127.0.0.1 --port 8000
The first command makes the UI accessible through your PC's IP on a trusted network.
The second serves POST /v1/systemone; image paths resolve under the selected image root.
The interfaces serialize model requests. They do not provide a public multi-tenant service.
The historical Korean interface and 2048 experiments remain in apps/playground/.
See its instructions.
Data, evidence and attribution
- Training and reproduction: staged training, provenance and limits.
- Dataset publication: 15 historical corpus configurations, original splits, image hashes and source-specific licenses. Configurations overlap; do not concatenate them.
- Research evidence: successful and failed experiments.
- Release history and future research.
Code and adaptation weights: Apache-2.0. Qwen retains its upstream Apache-2.0 attribution. Dataset licenses differ by source. LAYA and JEV inspired typed decision interfaces; their weights and code are not included, and no affiliation or equivalent performance is claimed. See NOTICE and source licenses.
Support
If QEV is useful, star the repository or visit
GitHub Sponsors. qev star --open opens
the repository for you to choose; using QEV does not require a star or donation.
Metadata
Release files for qev 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| qev-0.2.0.tar.gz | 171.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| qev-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 357.1 kB
Release files / qev-0.2.0.tar.gz
| Download URL | qev-0.2.0.tar.gz |
|---|---|
| Size | 171.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3f571a0e49407c696fc34e8d6ab0f7b4ea3250231974989b9e4a6c1586c5d3ba
|
|
BLAKE2b-256 checksum How to use checksums |
e78196d3e3537f6e107312218c23925b9c309c7f0c8c258a8cce41dd75e29e86
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / qev-0.2.0-py3-none-any.whl
| Download URL | qev-0.2.0-py3-none-any.whl |
|---|---|
| Size | 185.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b8a769ec18edce91f9699373df058df92c2297461ec49b327382cbc21dd31ef8
|
|
BLAKE2b-256 checksum How to use checksums |
6370986989e7a36147f0c69183afc461a4c0acf23f6651f1ad17b376f7d35cf3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log