Skip to main content

jev-checker

A Python semantic checker that adds Jev powered review diagnostics alongside Ruff and Pyrefly. Install from PyPI:

python -m pip install jev-checker

Quick start

Jev loads a .env file from the current directory or a parent directory. Put one provider key in that file; shell environment variables take precedence:

# .env
TYPESAFE_API_KEY=...
# Or use OPENROUTER_API_KEY=... and pass --provider openrouter

# Run with TypeSafe
jev check .

# Run with OpenRouter
jev check . --provider openrouter

The jev-checker check executable is an alias for jev check.

Check committed changes against the merge base of a branch:

jev check . --diff origin/main

--diff reviews commits only. It does not include uncommitted working tree changes. The source excerpts selected for review are sent to the chosen provider after credential pattern redaction. Redaction cannot identify every secret; review your provider's data policy before using the checker on sensitive code.

Diagnostics

The default output resembles other checkers:

src/auth.py:42:5: error: a protected operation may lack authorization [JEV201]

1 error, 0 warnings, 0 infos

error fails the check; warnings and infos do not unless --warnings-as-errors is set. Configure rule levels in pyproject.toml:

[tool.jev]
provider = "typesafe"
select = ["JEV1", "JEV2", "JEV3", "JEV4", "JEV5", "JEV6", "JEV7"]

[tool.jev.levels]
JEV305 = "warning"
JEV601 = "error"
JEV704 = "info"
JEV503 = "off"

The checker keeps diagnostic level, Jev probability, confidence, and impact score as separate values. Use --output-format full for evidence details, --show to print every screening score and follow-up decision, or json for a machine readable report including screening scores. Supported formats are concise, full, json, and github. Exit codes are 0 for no errors, 1 for blocking diagnostics, and 2 when the requested analysis cannot be completed. --exit-zero affects only code 1. --fail-on-unresolved also returns 2 for candidates without a reliable conclusion; --exit-zero does not override that explicit policy. JSON reports also distinguish supported, dismissed, context-limited, and incomplete candidates while omitting raw source excerpts.

Built-in rules

  • JEV1 correctness, JEV2 security, JEV3 reliability, JEV4 compatibility
  • JEV5 test evidence, JEV6 performance, JEV7 observability

Security, correctness, reliability, and compatibility rules default to errors. Test, performance, and observability rules default to warnings. Use --select, --ignore, and --rule-level RULE=LEVEL to adjust the run.

Reading --show

Each path:start-end heading is a source region. The score beside each rule is the model's screening probability from 0 to 1. * means the score reached screen-threshold (default 0.70) and was queued for follow-up; it is a signal, not a confirmed finding or a warning by itself. A diagnostic is emitted only after follow-up finds sufficiently confident, reachable evidence. These scores are model judgments, not calibrated guarantees.

Rule IDs

ID What it screens for
JEV101 A condition or branch handles the wrong cases.
JEV102 A value is sent to the wrong field, argument, or recipient.
JEV103 A state update breaks an established invariant.
JEV104 An accepted boundary input produces an incorrect result.
JEV105 An operation uses a dependent result before it is ready.
JEV201 A protected operation lacks required authorization.
JEV202 A user can access another user's or tenant's data or operations.
JEV203 Untrusted input reaches an interpreter or execution mechanism unsafely.
JEV204 Untrusted input escapes a filesystem or network boundary.
JEV205 Sensitive data reaches an unauthorized output or recipient.
JEV206 A default configuration leaves a required safeguard disabled.
JEV301 A failure path leaks an acquired resource.
JEV302 Concurrent work can update shared state inconsistently.
JEV303 A wait or resource acquisition can stall progress indefinitely.
JEV304 Cancellation leaves work or external effects inconsistent.
JEV305 Retrying after failure can duplicate a non-idempotent effect.
JEV401 A public contract conflicts with supplied consumers.
JEV402 A persisted or exchanged format conflicts with supplied readers.
JEV403 Exception or exit-code behavior conflicts with supplied consumers.
JEV404 Code needs a capability missing from a declared Python version or platform.
JEV501 Important behavior lacks targeted related-test evidence.
JEV502 An important failure or recovery path lacks test evidence.
JEV503 An authorization boundary lacks an access-denial test.
JEV504 An important boundary input lacks targeted test evidence.
JEV505 A test assertion may pass even when the relevant behavior is wrong.
JEV601 A blocking operation may run on an asynchronous event loop.
JEV602 Repeated external calls may have an equivalent batch alternative.
JEV603 Memory use may grow without a bound on input size.
JEV604 An invariant expensive calculation may be repeated unnecessarily.
JEV701 An important failure may end without an observable signal.
JEV702 A failed operation may be reported as successful.
JEV703 Distributed work may lose an identifier needed to connect its events.
JEV704 A diagnostic may contradict the operation's actual result.

The prefixes group rules: JEV1 correctness, JEV2 security, JEV3 reliability, JEV4 compatibility, JEV5 test evidence, JEV6 performance, and JEV7 observability. Their default levels are shown under Built-in rules.

Profiles and follow-up decisions

Source reviews screen each named function independently, including methods, async functions, and nested functions. Executable code outside functions is also screened. Every focused request includes the complete Python file as context, including its original line breaks. Localization selects evidence inside that focus; verification also receives the complete file. Candidates for the same rule in different functions are investigated independently. --show displays the qualified function names alongside screening scores and follow-up decisions. Related tests are selected by module imports or exact test-file naming conventions and are supplied in full. Source profiles also receive the full file. Credential literals remain redacted while preserving their positions.

The checker imposes no byte or line cap on this context. The provider enforces its own model limits: Jev currently documents 32k tokens for state plus the longest question and 64k tokens for the entire request (model limits). Bytes are not tokens. Rejected requests produce an incomplete review and exit code 2; source is never silently truncated to make a request fit. External files and contracts are only available when explicitly supplied through the review inputs.

File profiles show the file category and its confidence, then review priority and that score's confidence. Review priority runs from 0 (routine) to 3 (specialist or immediate review); it is not a diagnostic severity.

Follow-up decisions list each rule selected for closer review. p= repeats its screening probability. supported (diagnostic_emitted) means a diagnostic was produced; dismissed means follow-up did not confirm an issue; insufficient_context means confidence or evidence was too low; and incomplete means the follow-up could not finish. The reason in parentheses gives more detail, for example low_location_confidence means Jev could not pin the concern to a source region confidently enough to emit a diagnostic.

Checker pipeline

Run each tool as a separate step so its result stays visible:

ruff check .
ruff format --check .
pyrefly check
jev check . --diff origin/main --output-format json > jev-report.json

Provide the diff base in the CI checkout. Keep provider credentials in the CI secret store and run authenticated reviews only in trusted jobs. jev check . --dry-run reports file and region counts without using credentials or the network. There is no default cap on API attempts, evidence follow-ups, or run duration; all selected units and candidates are reviewed, with at most three concurrent requests. Use --max-requests N, --max-followups N, or --max-duration SECONDS to impose an explicit cap. A cap or deadline that interrupts review produces an incomplete result, not a clean check. The local SQLite cache fingerprints the complete file content. Unchanged files reuse judgments; any content edit automatically rechecks all review stages for that file, including functions whose own bodies did not change, because their contracts and callers may have changed. Other unchanged files keep their cache. Changes to the diff's before version or a supplied related test also invalidate the affected file's review. A timestamp change alone does not invalidate identical content.

The cache expires after 24 hours by default, including version-family aliases such as typesafe/jev-1.13. Explicit latest, auto, and preview aliases bypass it. For reproducible evaluations, use a provider-supported dated model identity and record the returned model. Normal use requires no --no-cache; that flag only forces a fresh run of unchanged content. The provider controls request pricing and rate limits.

Development

python -m pip install -e '.[dev]'
ruff check .
ruff format --check .
pyrefly check
pytest
pytest --cov=jev_review --cov-report=term-missing

The tests use simulated providers and do not make paid API calls. See docs/judgment-catalog.md for rule prompts and evidence criteria, and docs/implementation-plan.md for the implementation checklist and remaining release work.

Unresolved reviews and evidence integrity

The default output now exposes suspicious code that could not be confirmed:

No confirmed diagnostics; review has unresolved or unfinished work.
0 errors, 0 warnings, 0 infos
0 remote requests, 2 cache hits; models: jev-fixture
Unresolved: 1 candidate(s); this is not a clean review.
  sample.py [JEV101]: low_location_confidence

Use the following command to make unresolved work fail a verification pipeline:

jev check src/ --provider openrouter --fail-on-unresolved

The equivalent configuration is fail-on-unresolved = true under [tool.jev]. Without that policy, a completed run with only unresolved candidates retains exit 0 for compatibility. Confirmed errors return 1; operational failures, exhausted budgets, and the explicit unresolved policy return 2.

Screening remains at 0.70, verification at 0.80, and location/mechanism confidence at 0.70 by default. Intermediate verification probabilities remain unresolved; a sufficiently negative judgment dismisses the concern. Impact confidence describes uncertainty about severity and is reported separately from defect confirmation. Optional reviewer routing failures preserve confirmed diagnostics.

Evidence is split into bounded spans with exact ranges and content hashes. Oversized or omitted evidence is recorded; localization no longer silently sends only a prefix. Containing definitions and called local contracts are collected without executing reviewed code. Broad source locations are marked as evidence anchors in full output.

JSON schema migration

Reports now use schema_version: 2. Existing diagnostic, screening, candidate, usage, and operational status fields remain available. Consumers must accept version 2 and inspect review_status and unresolved_count: operational completion alone does not mean every candidate was resolved. New traces expose judgments, thresholds, evidence metadata, and remote/cache provenance without raw source or credentials. Cached judgments are reprocessed with current policy; prompt/state changes invalidate their cache keys. token_usage_status distinguishes reported usage from unavailable or partial usage and runs with no remote requests; token totals sum only known remote usage.

Semantic evaluation

Ordinary pytest checks use offline providers. The opt-in suite contains 31 buggy/fixed pairs, separate tuning/held-out splits, and static-analysis comparison controls. It measures actual model detection separately from pipeline tests:

python -m tools.evaluate_semantics --dry-run --split all
python -m tools.evaluate_semantics --provider openrouter --split tuning \
  --case success_inversion --case negative_transfer --case missing_authorization \
  --case failed_replacement --repetitions 3 --max-total-requests 100

Live evaluations send the repository-owned fixtures to the selected provider and may incur charges. The request budget applies to the entire run, including retries. They run without cache unless --cache-replay is explicit. Unresolved bugs count as missed detections; unattempted cases remain visible. A targeted rule evaluation is a tuning aid and cannot satisfy the end-to-end release gates.

See evaluation instructions and the reliability implementation checklist for measured limitations and remaining release criteria.

Experimental direct review

From the repository checkout, evaluate a direct per-function runtime judgment:

uv run python -m tools.scan_direct_review benchmarks/realistic_review \
  --provider openrouter --output .jev/direct-review.json
uv run python -m tools.evaluate_semantics --strategy direct --split all \
  --repetitions 3 --max-total-requests 1000 --output .jev/direct-corpus.json

This experiment uses complete-file context and reports suggestions at probability 0.70 without a second defect-confirmation gate. Uncertain positions are function anchors. It always uses fresh judgments and is separate from jev check. Reserved recall was 38/60 by expected rule, or 44/60 after verifying alternative rule classifications, with 3/60 corrected runs receiving false alerts. Three large-module findings were reproduced locally. The production quality gate remains unmet; see the experiment report.

Manual semantic scenarios

The scenario corpus adds 24 buggy/fixed pairs with obvious and subtle defects, varying impact, and executable local witnesses. It covers permissions, tenant isolation, money, concurrency, retries, formats, performance, and observability. Scan an individual example or directory:

uv run jev-checker check benchmarks/semantic_review/buggy --provider openrouter --show

The catalog explains each expected defect and its corresponding corrected file. Use the evaluation runner's --corpus benchmarks/semantic_review --split all option for repeated measurements with neutral paths and negative controls.

Pipeline benchmark

benchmarks/pipeline_review adds 5 buggy/fixed module pairs forming one small document-processing system, each defect in a different rule family (JEV301, JEV402, JEV604, JEV702, JEV205). Offline witnesses run with ordinary pytest; live detection runs through the same tools.evaluate_semantics --corpus benchmarks/pipeline_review entry point.

Calibration benchmark

benchmarks/calibration_review targets implicit (undocumented) authorization requirements, the pattern that motivated lowering verification_threshold's default to 0.80. Use --cache-replay with a different --verification-threshold to reprocess cached judgments at a new threshold without new API calls.

Metadata

Release files for jev-checker 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jev-checker 0.2.0
File Size Uploaded
jev_checker-0.2.0.tar.gz 215.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jev-checker 0.2.0
File Interpreter ABI Platform
jev_checker-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 263.6 kB

Release files / jev_checker-0.2.0.tar.gz

Download URL jev_checker-0.2.0.tar.gz
Size 215.8 kB
Tags Source
SHA-256 checksum
How to use checksums
d35b5fe7fc2ad56603f3e8a27ebbd674eda5f22da31c669caf78bd6ac89edb87
BLAKE2b-256 checksum
How to use checksums
60653a6759aad7ac4aa2eb72aee20854ac6704b4042018b05791fbb89ce86abd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / jev_checker-0.2.0-py3-none-any.whl

Download URL jev_checker-0.2.0-py3-none-any.whl
Size 47.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e182f24f97af6e530d78115cd0c616122633ed47a49264f5b3da7174d9d8ba15
BLAKE2b-256 checksum
How to use checksums
a0c523dd29a2fe123fc948ec7d8ecc68e45800619bb5181defd932a0ff10a435
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page