jev-checker
A Python semantic checker that adds Jev powered review diagnostics alongside Ruff and Pyrefly. Install from PyPI:
python -m pip install jev-checker
Quick start
Jev loads a .env file from the current directory or a parent directory. Put one
provider key in that file; shell environment variables take precedence:
# .env
TYPESAFE_API_KEY=...
# Or use OPENROUTER_API_KEY=... and pass --provider openrouter
# Run with TypeSafe
jev check .
# Run with OpenRouter
jev check . --provider openrouter
The jev-checker check executable is an alias for jev check.
Check committed changes against the merge base of a branch:
jev check . --diff origin/main
--diff reviews commits only. It does not include uncommitted working tree changes.
The source excerpts selected for review are sent to the chosen provider after
credential pattern redaction. Redaction cannot identify every secret; review your
provider's data policy before using the checker on sensitive code.
Diagnostics
The default output resembles other checkers:
src/auth.py:42:5: error: a protected operation may lack authorization [JEV201]
1 error, 0 warnings, 0 infos
error fails the check; warnings and infos do not unless --warnings-as-errors
is set. Configure rule levels in pyproject.toml:
[tool.jev]
provider = "typesafe"
select = ["JEV1", "JEV2", "JEV3", "JEV4", "JEV5", "JEV6", "JEV7"]
[tool.jev.levels]
JEV305 = "warning"
JEV601 = "error"
JEV704 = "info"
JEV503 = "off"
The checker keeps diagnostic level, Jev probability, confidence, and impact score
as separate values. Use --output-format full for evidence details, --show to
print every screening score and follow-up decision, or json for a machine
readable report including screening scores. Supported formats are concise,
full, json, and github. Exit codes are 0 for no errors, 1 for blocking diagnostics, and 2 when
the requested analysis cannot be completed. --exit-zero affects only code 1.
--fail-on-unresolved also returns 2 for candidates without a reliable conclusion;
--exit-zero does not override that explicit policy.
JSON reports also distinguish supported, dismissed, context-limited, and incomplete
candidates while omitting raw source excerpts.
Built-in rules
JEV1correctness,JEV2security,JEV3reliability,JEV4compatibilityJEV5test evidence,JEV6performance,JEV7observability
Security, correctness, reliability, and compatibility rules default to errors.
Test, performance, and observability rules default to warnings. Use --select,
--ignore, and --rule-level RULE=LEVEL to adjust the run.
Reading --show
Each path:start-end heading is a source region. The score beside each rule is
the model's screening probability from 0 to 1. * means the score reached
screen-threshold (default 0.70) and was queued for follow-up; it is a signal,
not a confirmed finding or a warning by itself. A diagnostic is emitted only
after follow-up finds sufficiently confident, reachable evidence. These scores
are model judgments, not calibrated guarantees.
Rule IDs
| ID | What it screens for |
|---|---|
| JEV101 | A condition or branch handles the wrong cases. |
| JEV102 | A value is sent to the wrong field, argument, or recipient. |
| JEV103 | A state update breaks an established invariant. |
| JEV104 | An accepted boundary input produces an incorrect result. |
| JEV105 | An operation uses a dependent result before it is ready. |
| JEV201 | A protected operation lacks required authorization. |
| JEV202 | A user can access another user's or tenant's data or operations. |
| JEV203 | Untrusted input reaches an interpreter or execution mechanism unsafely. |
| JEV204 | Untrusted input escapes a filesystem or network boundary. |
| JEV205 | Sensitive data reaches an unauthorized output or recipient. |
| JEV206 | A default configuration leaves a required safeguard disabled. |
| JEV301 | A failure path leaks an acquired resource. |
| JEV302 | Concurrent work can update shared state inconsistently. |
| JEV303 | A wait or resource acquisition can stall progress indefinitely. |
| JEV304 | Cancellation leaves work or external effects inconsistent. |
| JEV305 | Retrying after failure can duplicate a non-idempotent effect. |
| JEV401 | A public contract conflicts with supplied consumers. |
| JEV402 | A persisted or exchanged format conflicts with supplied readers. |
| JEV403 | Exception or exit-code behavior conflicts with supplied consumers. |
| JEV404 | Code needs a capability missing from a declared Python version or platform. |
| JEV501 | Important behavior lacks targeted related-test evidence. |
| JEV502 | An important failure or recovery path lacks test evidence. |
| JEV503 | An authorization boundary lacks an access-denial test. |
| JEV504 | An important boundary input lacks targeted test evidence. |
| JEV505 | A test assertion may pass even when the relevant behavior is wrong. |
| JEV601 | A blocking operation may run on an asynchronous event loop. |
| JEV602 | Repeated external calls may have an equivalent batch alternative. |
| JEV603 | Memory use may grow without a bound on input size. |
| JEV604 | An invariant expensive calculation may be repeated unnecessarily. |
| JEV701 | An important failure may end without an observable signal. |
| JEV702 | A failed operation may be reported as successful. |
| JEV703 | Distributed work may lose an identifier needed to connect its events. |
| JEV704 | A diagnostic may contradict the operation's actual result. |
The prefixes group rules: JEV1 correctness, JEV2 security, JEV3
reliability, JEV4 compatibility, JEV5 test evidence, JEV6 performance,
and JEV7 observability. Their default levels are shown under Built-in rules.
Profiles and follow-up decisions
Source reviews screen each named function independently, including methods,
async functions, and nested functions. Executable code outside functions is also
screened. Every focused request includes the complete Python file as context,
including its original line breaks. Localization selects evidence inside that
focus; verification also receives the complete file. Candidates for the same rule
in different functions are investigated independently. --show displays the
qualified function names alongside screening scores and follow-up decisions.
Related tests are selected by module imports or exact
test-file naming conventions and are supplied in full. Source profiles also receive
the full file. Credential literals remain redacted while preserving their positions.
The checker imposes no byte or line cap on this context. The provider enforces its own model limits: Jev currently documents 32k tokens for state plus the longest question and 64k tokens for the entire request (model limits). Bytes are not tokens. Rejected requests produce an incomplete review and exit code 2; source is never silently truncated to make a request fit. External files and contracts are only available when explicitly supplied through the review inputs.
File profiles show the file category and its confidence, then review priority
and that score's confidence. Review priority runs from 0 (routine) to 3
(specialist or immediate review); it is not a diagnostic severity.
Follow-up decisions list each rule selected for closer review. p= repeats
its screening probability. supported (diagnostic_emitted) means a diagnostic
was produced; dismissed means follow-up did not confirm an issue;
insufficient_context means confidence or evidence was too low; and
incomplete means the follow-up could not finish. The reason in parentheses
gives more detail, for example low_location_confidence means Jev could not
pin the concern to a source region confidently enough to emit a diagnostic.
Checker pipeline
Run each tool as a separate step so its result stays visible:
ruff check .
ruff format --check .
pyrefly check
jev check . --diff origin/main --output-format json > jev-report.json
Provide the diff base in the CI checkout. Keep provider credentials in the CI
secret store and run authenticated reviews only in trusted jobs. jev check . --dry-run reports file and region counts without using credentials or the network.
There is no default cap on API attempts, evidence follow-ups, or run duration;
all selected units and candidates are reviewed, with at most three concurrent
requests. Use --max-requests N, --max-followups N, or --max-duration SECONDS
to impose an explicit cap. A cap or deadline that interrupts review produces an
incomplete result, not a clean check.
The local SQLite cache fingerprints the complete
file content. Unchanged files reuse judgments; any content edit automatically
rechecks all review stages for that file, including functions whose own bodies
did not change, because their contracts and callers may have changed. Other
unchanged files keep their cache. Changes to the diff's
before version or a supplied related test also invalidate the affected file's
review. A timestamp change alone does not invalidate identical content.
The cache expires after 24 hours by default, including version-family aliases
such as typesafe/jev-1.13. Explicit
latest, auto, and preview aliases bypass it. For reproducible evaluations,
use a provider-supported dated model identity and record the returned model.
Normal use requires no --no-cache; that flag only forces a fresh run of unchanged
content. The provider controls request pricing and rate limits.
Development
python -m pip install -e '.[dev]'
ruff check .
ruff format --check .
pyrefly check
pytest
pytest --cov=jev_review --cov-report=term-missing
The tests use simulated providers and do not make paid API calls. See
docs/judgment-catalog.md for rule prompts and evidence
criteria, and docs/implementation-plan.md for the
implementation checklist and remaining release work.
Unresolved reviews and evidence integrity
The default output now exposes suspicious code that could not be confirmed:
No confirmed diagnostics; review has unresolved or unfinished work.
0 errors, 0 warnings, 0 infos
0 remote requests, 2 cache hits; models: jev-fixture
Unresolved: 1 candidate(s); this is not a clean review.
sample.py [JEV101]: low_location_confidence
Use the following command to make unresolved work fail a verification pipeline:
jev check src/ --provider openrouter --fail-on-unresolved
The equivalent configuration is fail-on-unresolved = true under [tool.jev].
Without that policy, a completed run with only unresolved candidates retains exit
0 for compatibility. Confirmed errors return 1; operational failures, exhausted
budgets, and the explicit unresolved policy return 2.
Screening remains at 0.70, verification at 0.80, and location/mechanism confidence at 0.70 by default. Intermediate verification probabilities remain unresolved; a sufficiently negative judgment dismisses the concern. Impact confidence describes uncertainty about severity and is reported separately from defect confirmation. Optional reviewer routing failures preserve confirmed diagnostics.
Evidence is split into bounded spans with exact ranges and content hashes. Oversized or omitted evidence is recorded; localization no longer silently sends only a prefix. Containing definitions and called local contracts are collected without executing reviewed code. Broad source locations are marked as evidence anchors in full output.
JSON schema migration
Reports now use schema_version: 2. Existing diagnostic, screening, candidate,
usage, and operational status fields remain available. Consumers must accept
version 2 and inspect review_status and unresolved_count: operational completion
alone does not mean every candidate was resolved. New traces expose judgments,
thresholds, evidence metadata, and remote/cache provenance without raw source or
credentials. Cached judgments are reprocessed with current policy; prompt/state
changes invalidate their cache keys. token_usage_status distinguishes reported
usage from unavailable or partial usage and runs with no remote requests; token totals
sum only known remote usage.
Semantic evaluation
Ordinary pytest checks use offline providers. The opt-in suite contains 31 buggy/fixed pairs, separate tuning/held-out splits, and static-analysis comparison controls. It measures actual model detection separately from pipeline tests:
python -m tools.evaluate_semantics --dry-run --split all
python -m tools.evaluate_semantics --provider openrouter --split tuning \
--case success_inversion --case negative_transfer --case missing_authorization \
--case failed_replacement --repetitions 3 --max-total-requests 100
Live evaluations send the repository-owned fixtures to the selected provider and
may incur charges. The request budget applies to the entire run, including retries.
They run without cache unless --cache-replay is explicit. Unresolved bugs count
as missed detections; unattempted cases remain visible. A targeted rule evaluation
is a tuning aid and cannot satisfy the end-to-end release gates.
See evaluation instructions and the reliability implementation checklist for measured limitations and remaining release criteria.
Experimental direct review
From the repository checkout, evaluate a direct per-function runtime judgment:
uv run python -m tools.scan_direct_review benchmarks/realistic_review \
--provider openrouter --output .jev/direct-review.json
uv run python -m tools.evaluate_semantics --strategy direct --split all \
--repetitions 3 --max-total-requests 1000 --output .jev/direct-corpus.json
This experiment uses complete-file context and reports suggestions at probability
0.70 without a second defect-confirmation gate. Uncertain positions are function
anchors. It always uses fresh judgments and is separate from jev check.
Reserved recall was 38/60 by expected rule, or 44/60 after verifying alternative
rule classifications, with 3/60 corrected runs receiving false alerts. Three
large-module findings were reproduced locally. The production quality gate
remains unmet; see the experiment report.
Manual semantic scenarios
The scenario corpus adds 24 buggy/fixed pairs with obvious and subtle defects, varying impact, and executable local witnesses. It covers permissions, tenant isolation, money, concurrency, retries, formats, performance, and observability. Scan an individual example or directory:
uv run jev-checker check benchmarks/semantic_review/buggy --provider openrouter --show
The catalog explains each expected defect and its corresponding corrected file.
Use the evaluation runner's --corpus benchmarks/semantic_review --split all
option for repeated measurements with neutral paths and negative controls.
Pipeline benchmark
benchmarks/pipeline_review adds 5
buggy/fixed module pairs forming one small document-processing system, each
defect in a different rule family (JEV301, JEV402, JEV604, JEV702, JEV205).
Offline witnesses run with ordinary pytest; live detection runs through the
same tools.evaluate_semantics --corpus benchmarks/pipeline_review entry point.
Calibration benchmark
benchmarks/calibration_review
targets implicit (undocumented) authorization requirements, the pattern that
motivated lowering verification_threshold's default to 0.80. Use
--cache-replay with a different --verification-threshold to reprocess
cached judgments at a new threshold without new API calls.
Metadata
Release files for jev-checker 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| jev_checker-0.2.0.tar.gz | 215.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| jev_checker-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 263.6 kB
Release files / jev_checker-0.2.0.tar.gz
| Download URL | jev_checker-0.2.0.tar.gz |
|---|---|
| Size | 215.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d35b5fe7fc2ad56603f3e8a27ebbd674eda5f22da31c669caf78bd6ac89edb87
|
|
BLAKE2b-256 checksum How to use checksums |
60653a6759aad7ac4aa2eb72aee20854ac6704b4042018b05791fbb89ce86abd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / jev_checker-0.2.0-py3-none-any.whl
| Download URL | jev_checker-0.2.0-py3-none-any.whl |
|---|---|
| Size | 47.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e182f24f97af6e530d78115cd0c616122633ed47a49264f5b3da7174d9d8ba15
|
|
BLAKE2b-256 checksum How to use checksums |
a0c523dd29a2fe123fc948ec7d8ecc68e45800619bb5181defd932a0ff10a435
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log