Skip to main content

jev-tools

Let a tiny decision model make the cheap, frequent judgments, so Claude doesn't have to.

A Claude Code plugin for the terminal and the desktop app. Does this edit break a project rule? Which skill fits this prompt? Which file matters? Does this diff need a careful review?

PyPI License MIT Python 3.10+ stdlib only Claude Code plugin shadow-first

Quick start · Choose a backend · What's inside · Privacy · Troubleshooting · Reference


jev-tools asks OpenJev (Codiv "System One") typed questions (yes/no, pick-one, score) and gets calibrated probabilities back instead of generated text. There is nothing to parse and nothing to hallucinate; your own code applies the thresholds. It runs against the hosted API, a model on your own machine, or in an offline pattern-only mode.

Write-up with the numbers, failures and privacy notes: I built a Claude Code plugin around a model that only answers yes/no · feedback welcome in issue #1.

🚀 Quick start

You need: Python 3.10+ as the command python (the hooks call exactly that), the claude CLI, and uv or pipx.

# 1. install the setup tool (once)
uv tool install jev-tools-setup          # or: pipx install jev-tools-setup

# 2. optional, read-only: is this machine ready?
jev-tools-setup check

# 3. install the plugin and choose how Jev runs
jev-tools-setup

Setup installs the skills, scripts and hooks through Claude Code (so terminal and desktop share one install), asks which backend you want, validates it, and finishes by writing ~/.jev-tools/LOG.md. Restart Claude Code, then run the status skill to confirm everything is wired.

Prefer a one-shot run without installing anything? uvx jev-tools-setup check and uvx jev-tools-setup work too. To run the latest unreleased code straight from GitHub, use uvx --from git+https://github.com/Rcidshacker/jev-tools jev-tools-setup.

No uv or pipx? Install the plugin by hand

Inside Claude Code:

/plugin marketplace add Rcidshacker/jev-tools
/plugin install jev-tools@jev-tools

Then give it a key (or a local server) yourself, see Reference. Try it for one session without installing: claude --plugin-dir /path/to/jev-tools.

🔀 Choose how Jev runs

api local offline
What it is Hosted OpenJev at api.codiv.ai A model on your machine No model, pattern fallbacks
Best for Best accuracy, zero setup beyond a key Privacy, no key, no quota Plugin installed but nothing sent anywhere
Needs A free Codiv key (100M input tokens) Verdict/Laya: any CPU or GPU. Full OpenJev: NVIDIA Blackwell-class GPU + Docker Nothing
Leaves your machine Diffs, prompts, rules (see Privacy) Nothing (weights download once from Hugging Face) Nothing
Model quality Full OpenJev (all measurements below) Verdict and Laya are small: they only judge what they proved able to (see measured) n/a
api: paste a key, it is checked and stored safely

Setup asks for your key in a hidden prompt, validates it with one live call, and stores it in ~/.jev-tools/credentials with owner-only permissions (chmod 600, or icacls on Windows). A rejected key saves nothing. The key is never passed on a command line and never logged.

local: Verdict 151M, Laya 421M, or the full OpenJev model

Setup asks which model:

Model Size Context Options Notes
verdict-1.4 151M 512 tokens up to 24 Smallest. Ignores a yes/no question's criteria. Runs on CPU or any GPU
laya-1.0 421M 1,024 tokens about 20 Heavier. Runs on CPU or any GPU
openjev-latest 26B 65,536 tokens 255 Needs an NVIDIA Blackwell-class GPU (NVFP4) and Docker (self-hosting guide)

For Verdict/Laya, setup shows the plan and the download size, waits for your yes, then clones OpenJev into ~/.jev-tools, creates a venv and installs the model's extra (PyTorch comes with it and can be several GB; weights are about 0.6 GB / 1.7 GB as fp32). Start the server in its own terminal and keep it open:

jev-tools-setup serve        # add --device cpu to force the CPU

What changes with a small model. Each question is read with its own copy of the state, cut from the end at the model's window (512 / 1,024 tokens). jev-tools trims long inputs to fit, caps option lists (23 for Verdict, 19 for Laya) and declines a question it cannot shrink. Local servers get one request at a time: they are CPU/GPU bound, and parallel requests only queue and time out.

The plugin only asks a small model what it proved able to judge (live runs below). Everything else automatically uses the same pattern fallbacks as offline, and status and LOG.md say so:

Asks the model Uses pattern fallbacks
Verdict skill hook rule hook, review-precheck, find-files, browser-nav
Laya browser-nav skill hook (keywords), rule hook, review-precheck, find-files
offline: "offline" means no model, not no internet

Nothing is sent anywhere. Deterministic fallbacks still run, and each says what it cannot do:

Piece Offline behaviour
review-precheck Regex checks for secrets, dependency files, auth keywords, schema/migration changes, test weakening and swallowed errors. Says "fast" only for a small (30 lines or fewer) clean diff
Rule hook Blocks only a literal token that a "never / don't / avoid X" rule forbids (e.g. console.log). No judgment
Skill hook Picks by keyword match
find-files Ranks by keyword count (in our measurements the model was no better, see below)
browser-nav Suggests the element whose name matches the goal, always marked unsure
rule-calibrate Unavailable: it needs a model, and says so

These are patterns, not understanding. They miss anything phrased unusually, so keep offline in shadow.

🧰 What's inside

Piece Kind What it does
Rule hook rule_enforcer.py PreToolUse on Edit|Write|MultiEdit Reads your CLAUDE.md / AGENTS.md rules, asks whether the pending edit breaks one, re-checks any hit with a stricter question, and in active mode blocks the write with the rule it broke
Skill hook skill_picker.py UserPromptSubmit (opt-in) Picks the one installed skill that fits your prompt, confirms it, and in active mode adds a one-line pointer to the context
find-files skill Keyword candidates, then two-stage scoring, prints the most relevant files. A hint, not an oracle
browser-nav skill Picks each next click from the page's interactive elements; Claude executes and judges pass/fail
review-precheck skill Seven yes/no policy questions on a git diff decide fast pass vs full review
rule-calibrate skill Replays recent commits against your rules: which are decisive, noisy, weak or quiet, before you enforce anything
status skill One screen: backend, key found (by name), mode, hook state, recent events and errors
report skill Writes the plain-English LOG.md

⚙️ How it works

flowchart LR
    A["Claude Code event<br/>edit · prompt · skill"] --> B["jev-tools script"]
    B --> C{"backend"}
    C -->|api| D["api.codiv.ai"]
    C -->|local| E["Verdict · Laya · OpenJev<br/>127.0.0.1:8080"]
    C -->|offline| F["keyword and pattern<br/>fallbacks"]
    D --> G["typed answers<br/>yes/no · pick · score"]
    E --> G
    G --> H["thresholds in plain code"]
    F --> H
    H --> I{"mode"}
    I -->|shadow| J["log only"]
    I -->|active| K["block or inject"]

Each piece turns a fuzzy judgment into typed questions, sends them with the relevant text as state to POST /v1/systemone, and applies thresholds in ordinary code. Every failure path is explicit: hooks fail open (an outage never blocks your edit or prompt) and the review pre-check fails safe (an outage routes to full review). Details and the reasoning behind each threshold: HOW_IT_WORKS.md.

Modes

Mode Rule hook Skill hook
shadow (default) logs the verdict, never blocks logs the pick, injects nothing
active blocks an edit flagged at ≥ 0.80 and confirmed at ≥ 0.70 injects Jev skill pick: <name> (if enabled)
off does nothing, makes no call does nothing

Recommended path: run in shadow for a week, read LOG.md, run rule-calibrate on a project, then switch that setup to active. Set the mode with the plugin's mode option or JEV_MODE (the variable wins).

🔒 Privacy: what leaves your machine

With the api backend this plugin sends text to a third party (api.codiv.ai). local and offline send nothing. Read this before enabling api on any project.

Piece What is sent (api only)
Rule hook (every Edit/Write) file name, the unified diff (up to 8,000 characters) and the text of your project rules
Skill hook (opt-in, every prompt) your prompt, plus the names and descriptions of your installed skills
find-files the query and short excerpts of candidate files
review-precheck / rule-calibrate the git diff (up to 30k characters) / recent commit hunks
browser-nav the goal, the URL and the names of the page's interactive elements
  • Shadow mode still sends. The hook needs the model's answer to log it. Only mode: off sends nothing.
  • A local seatbelt runs first. Before sending, these become [REDACTED]: private-key blocks, vendor-style keys (sk-/sk_live_, AWS, Google, GitHub, Slack, npm), JWTs, Bearer tokens, credentials in connection strings, assignments to names containing password/secret/token/api key, and long hex or base64-looking strings. Files named like secrets (.env*, *.pem, *.key, id_rsa*, credentials*, secrets* …) are never read into a request. It is a pattern match: a bare token in an unusual shape, names, emails, customer or employee records and internal business data are not caught, and about 0.2% of ordinary code lines that mention token/secret get partly redacted. Risk reduction, not a guarantee.
  • Retention is unknown to us. Codiv's public API docs say nothing about how request data is stored or used. Check their terms before sending anything you would not paste into a public forum.
  • Nothing sensitive is stored locally. The decision log holds verdicts, probabilities, latency and token counts, never file contents and never the key.

Turn it off for one project (regulated, customer, employee or client data) in that project's .claude/settings.local.json:

{ "env": { "JEV_MODE": "off" } }

🩺 When something looks wrong

jev-tools-setup always ends by writing ~/.jev-tools/LOG.md, even when setup fails. Refresh it any time with jev-tools-setup report, or run the report skill inside Claude Code. It merges every decision log (the plugin data directory and ~/.jev-tools) into plain English:

  • a one-line health verdict: ✅ healthy · ⚠️ notes · ❌ problems
  • your setup: backend, model, mode, which variable the key came from (never the key), which logs were read
  • problems, each with a fix
  • what happened: event counts, median latency, rule checks flagged
  • your setup history and the last 40 events, one sentence each

It contains no API key, file contents or prompts; home paths show as ~ and key-shaped strings are redacted, so it is safe to attach to an issue.

Symptom Likely cause Fix
Hooks do nothing and nothing is logged JEV_MODE=off (the variable beats the saved mode), or Claude Code was not restarted Unset JEV_MODE, restart Claude Code
status says the key is MISSING No key in the environment, plugin option or credentials file jev-tools-setup, choose api
401 / 403 in LOG.md Key rejected New key at codiv.ai/dashboard, re-run setup
429 in LOG.md Quota or rate limit (quota errors are not retried) Wait, or check usage on the dashboard
python not found, or opens the Microsoft Store The hooks call the literal command python Install Python 3.10+; jev-tools-setup check tests it
Local server not answering It is not running, or still loading weights jev-tools-setup serve
Bad edits are never blocked You are in shadow (the default) Calibrate, then switch to active
review-precheck always says "full" Offline backend, a small model cutting the diff, or an outage LOG.md says which

Hooks fail open, so a broken install looks identical to a working one until you look at the log. That is exactly what LOG.md and the status skill are for. To remove everything: jev-tools-setup uninstall deletes ~/.jev-tools (key, config, log); remove the plugin with /plugin uninstall jev-tools@jev-tools.

📚 Reference

Giving it a key by hand

Any one of these, never in a repo. Lookup order: environment, plugin option, then the setup credentials file.

  • Environment variable: a user-level OPENJEV_API_KEY. Windows: setx OPENJEV_API_KEY "sk-codiv-..." then restart Claude Code. macOS/Linux: export it in your shell profile.
  • Plugin option: Claude Code prompts for the api_key setting and stores it as sensitive.
  • jev-tools-setup: writes the owner-only credentials file for you.

Do not put the key in settings.json, CLAUDE.md or any file in a repo. The client only sends it to api.codiv.ai (or loopback), refuses redirects and never logs it. python scripts/jevlib.py makes one live call and prints OK.

Environment variables
Variable Default Meaning
OPENJEV_API_KEY none API key (also TYPESAFE_API_KEY, or the plugin's api_key option)
OPENJEV_BASE_URL https://api.codiv.ai Must be api.codiv.ai or loopback
JEV_MODE shadow shadow, active or off; overrides the saved mode
JEV_SKILL_PICKER off 1 enables the skill hook (also the skill_picker option)
JEV_THRESHOLD 0.80 Rule-violation probability that can block
JEV_CONFIRM 0.70 Second-look probability required to block
JEV_SKILL_MIN 0.60 Minimum probability to inject a skill pick
JEV_PRECHECK_MIN 0.25 Yes-probability that flags a diff for full review
JEV_LOG plugin data dir Where decisions are appended (JSONL)
Files jev-tools writes
File Holds
~/.jev-tools/config.json backend, model, base_url (local only), mode
~/.jev-tools/credentials your API key, owner-only
~/.jev-tools/log.jsonl and the plugin data dir's log.jsonl one JSON line per decision: verdict, probability, latency, token counts
~/.jev-tools/LOG.md the plain-English report built from the logs
~/.jev-tools/openjev, ~/.jev-tools/venv local small-model install
What we measured (full OpenJev, small samples, live against Codiv)

All reproducible from the scripts; method and caveats in MEASUREMENTS.md.

Component Result
Rule enforcer 4 of 4 planted violations blocked, 4 of 4 clean edits allowed after the second-look check was added; also blocked and allowed correctly inside real headless Claude Code sessions
Skill picker 3 of 3 correct on a large real skill roster (two matches, one correct "none")
Review pre-check benign rename routed fast; a diff with a hardcoded key, swallowed exception and emptied tests routed full with the right flags
Browser navigator 5 of 5 steps correct on a synthetic login flow (never run against a real browser)
File discovery no better than plain keyword counting on 8 labelled queries (top-3 hits 4 to 6 of 8 vs 4 of 8); repeat runs differ by up to 2
Cost and speed about 1 s per prompt or edit, 2 s when a violation is confirmed; about 5k input tokens per edit at 20 rules

Not measured on real projects: the offline fallbacks. The small local models are measured in the next section.

Measured on real small models (Verdict 151M and Laya 421M, CPU, one Windows laptop)

Run against live local servers installed by jev-tools-setup, with the real scripts. Tiny samples (2 to 8 cases each), so read them as "can it do this at all", not as accuracy figures.

Task Verdict 151M Laya 421M
Plain topical choice (invoice / ticket / incident / other) 4 of 4 4 of 4
Skill picking (3 cases) 3 of 3 through the model picks were right, but their scores sat under the 0.6 bar tuned on OpenJev (0.47; the second was vetoed), so nothing was injected; the plugin now uses its keyword pick for Laya
browser-nav (2 steps) 1 of 2 right, both under the confidence floor 2 of 2 at 0.88 and 0.94
Rule check, violation vs clean edit (6 cases) scores flat near 0.65 for everything: mean gap about 0, 3 of 6 at chance violations scored 0.39 to 0.55 vs 0.16 to 0.29 clean: a gap, but under the 0.8 block line
review-precheck flagged all 7 policy questions even on a rename flagged 3 of 7 on a rename, 7 of 7 on the risky diff
find-files (8 labelled queries, top 3) 0 of 8, about 7 s per query 0 of 8, about 56 s per query
Latency, 16 questions about 0.55 s about 1.5 s

Plain keyword counting got 2 of 8 on the same queries, which is why the plugin does not ask either model to rank files. With the pattern fallbacks, on the same cases: the rule hook blocked the planted console.log edit and allowed all clean edits (a rule such as "never commit secrets" is matched by secret shape; rules with neither a literal token nor a secret topic are not caught), and review-precheck called the benign rename fast and the risky diff full with the right flags.

Takeaways: both models classify topics well; neither is a drop-in for the full model on code judgments; Laya is the better of the two but slower on CPU; run rule-calibrate before trusting any rule verdict from a small model.

Repository layout
.claude-plugin/plugin.json       plugin manifest and user settings (api_key, mode, skill_picker)
.claude-plugin/marketplace.json  single-plugin marketplace so /plugin marketplace add works
installer/jev_tools_cli.py       the jev-tools-setup command (setup, check, serve, report, uninstall)
installer/jev_report.py          builds LOG.md (also run by the report skill)
hooks/hooks.json                 the two hooks
scripts/                         jevlib.py (client) and one script per piece, review_policy.json
skills/<name>/SKILL.md           six skills
tests/test_all.py                offline suite against a local mock of /v1/systemone
docs/                            how it works, measurements, build log, Codiv API notes
pyproject.toml                   packages the installer as jev-tools-setup

🛠️ Development

python tests/test_all.py          # offline suite, no network, no key needed
python scripts/jevlib.py          # one live call, needs OPENJEV_API_KEY
uv build                          # builds the jev-tools-setup wheel and sdist
# releases: publishing a GitHub release runs .github/workflows/publish.yml (PyPI trusted publishing, no token)

The tests spin up a local server that mimics /v1/systemone, so they prove the logic and the wire format, not OpenJev's accuracy. Accuracy claims come only from the live runs recorded in the docs. See the CHANGELOG for what changed in each version.

🙏 Credits

Ideas and hard-won numbers borrowed, with thanks, from projects that got there first (exactly what was taken from each: BUILD_LOG.md): abide (rule compilation, calibration verdicts, second look) · hermes-jev-skills (confidence floor, two-stage retrieval) · jev-kit (shadow-first rollout) · jevgate (confirming before acting) · jev-skill-router (plugin layout, honest field report on skill routing). The small local models are Verdict by Heman10x and Laya by Nandakishor M / Convai Innovations, served by OpenJev.

Built with Claude Code. MIT licensed, see LICENSE.

Metadata

Release files for jev-tools-setup 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jev-tools-setup 0.3.1
File Size Uploaded
jev_tools_setup-0.3.1.tar.gz 46.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jev-tools-setup 0.3.1
File Interpreter ABI Platform
jev_tools_setup-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 72.4 kB

Release files / jev_tools_setup-0.3.1.tar.gz

Download URL jev_tools_setup-0.3.1.tar.gz
Size 46.3 kB
Tags Source
SHA-256 checksum
How to use checksums
6fb7187f076c1352b38b32a6f5d3ada1704bee55a6fe8bad426c3c5e02af1f76
BLAKE2b-256 checksum
How to use checksums
a45575ece3bffdd92fd7703ec0e0d9b02ef2980add260a8e065efc1508207bbc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / jev_tools_setup-0.3.1-py3-none-any.whl

Download URL jev_tools_setup-0.3.1-py3-none-any.whl
Size 26.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a9f33e85d8ae698a725ead39628b337d568a7468a730e441e6dc8b7a36cd0e79
BLAKE2b-256 checksum
How to use checksums
38d98fbdd76c724f2e2cd4c7fb3d16dea0ada281d9b806734b4256315c2a129f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page