Skip to main content

j-tools: semantic judgment for the shell

ci PyPI python MIT tests

j-tools

Semantic judgment for the shell. Ten small Unix tools that sort, pick, gate, watch, dedupe, tag, route and join text by meaning. Each does one decision job and pipes into the next, exactly like grep, sort, uniq and head. The difference: the criterion is a sentence, and the decision is made by Jev, TypeSafe's decision model, in about

240 ms for about a thousandth of a cent.
$ cat feedback.txt | jsort "angriest customer first" --with-score
0.950	WHY does it log me out every five minutes?? Absolutely unacceptable.
0.800	Cancel my subscription immediately and refund the last charge. Worst support ever.
0.700	Your sales rep promised a feature that does not exist. I feel lied to.
0.050	Love the new dashboard, thanks team!

$ tail -f app.log | jwatch "something a human should look at right now" --cooldown 300
2026-09-18T10:00:05Z ERROR api    payment-svc unreachable, 3 retries exhausted, order 88213 not charged

$ git diff | jgate "this change is safe to auto-merge" && git merge
$ cat inbox.txt | jroute "sales:a sales lead" "support:a support request" "spam:junk" --out-dir sorted

jsort demo: real output, recorded

Jev is not a text generator. It takes a state (text or JSON) and typed questions, and returns typed answers with calibrated probabilities. Nothing here ever asks a model to write anything. That is the whole design, and it is what makes these tools fast enough and cheap enough to sit in a pipe.

the three primitives: noul, choice, score

The ten tools

tool one decision job primitive try it
jgrep print lines that fit a description noul tail -f app.log | jgrep "a user is getting frustrated"
jsort sort lines by how well they fit score cat tasks.txt | jsort "most urgent" -n 10
jpick choose the single best line choice cat subject-lines.txt | jpick "most likely to get opened" --why
jgate exit status from a judgment; guards && noul cat report.txt | jgate "mentions a security incident" && alert
jwatch alert only on lines worth attention, live noul journalctl -f | jwatch "a service is failing repeatedly" --exec 'notify-send {}'
juniq drop lines that mean the same as an earlier one noul × pairs cat requests.txt | juniq "same underlying request" -c
jhead the N most relevant lines, in original order score cat thread.txt | jhead 5 "the key decisions made"
jtag add a label or score column, nothing dropped choice / score cat tickets.txt | jtag --labels "bug,feature,question"
jroute split a stream into bucket files choice cat inbox.txt | jroute "sales:a lead" "spam:junk" -o sorted
jmatch semantic join: for each line in A, the best in B choice jmatch invoices.txt payments.txt "same transaction"

Plus jtools (list, doctor). Every tool reads stdin or files, writes stdout, preserves input order unless its job is to reorder, streams (tail -f works), and shares one set of flags and exit codes. More pipelines in docs/recipes.md.

Real output from each tool (recorded, not typed; click to expand)
jroute jwatch
juniq jtag
jmatch jhead
jpick jgrep
jgate

Animated versions: docs/assets/demo-*.gif; asciinema casts: docs/demo/*.cast.

Install

uv tool install jev-tools                                 # all eleven commands
uv tool install git+https://github.com/adiel-hub/Jtools   # or straight from main

The suite is called j-tools; the package is jev-tools, after the model it asks. What you type day to day is the eleven commands above. (jtools on PyPI is an unrelated project.)

Python 3.11+. One dependency (httpx). Then one key, any of these:

backend key get one
TypeSafe (native) TYPESAFE_API_KEY console.typesafe.ai
OpenRouter OPENROUTER_API_KEY openrouter.ai/keys
Vercel AI Gateway AI_GATEWAY_API_KEY (vck_…) vercel.com/ai-gateway
your own System One gateway JEV_GATEWAY_URL + JEV_GATEWAY_API_KEY LiteLLM, a corporate proxy, a mock

Set it for this shell, or once and for good:

export AI_GATEWAY_API_KEY="vck_..."          # add it to ~/.zshrc or ~/.bashrc to keep it

mkdir -p ~/.config/jev                       # or a file every tool reads, no shell config at all
echo "vck_..." > ~/.config/jev/vercel.key
chmod 600 ~/.config/jev/vercel.key

The file is named after the backend: typesafe.key, openrouter.key, vercel.key, gateway.key (+ gateway.url). With several keys the first row of that table wins; force one with --api NAME or JEV_API.

Then check it — jtools doctor shows every backend it can see, names the one it picked, and makes one real call:

$ jtools doctor
jev-tools 0.1.1  python 3.12.4
config dir: ~/.config/jev   cache dir: ~/.cache/jev
  typesafe    no key     TYPESAFE_API_KEY
  openrouter  no key     OPENROUTER_API_KEY
  vercel      key found  AI_GATEWAY_API_KEY
  gateway     no key     JEV_GATEWAY_API_KEY

using vercel at https://ai-gateway.vercel.sh/v4/ai/evaluation-model with key vck_…9f2a, model typesafe-ai/jev
ok: one call in 412 ms, 283 input tokens, $0.0000119, answered by typesafe-ai/jev

With no key at all, every tool exits 3 and tells you the same four places to get one.

Shell completions for all eleven commands are in completions/:

source completions/j-tools.bash            # bash
cp completions/_jtools ~/.zfunc/ && echo 'fpath+=~/.zfunc' >> ~/.zshrc   # zsh

Where this sits

j-tools lx and other LLM shell wrappers grepai, ck and other embedding search
what the model does judges: yes/no, pick one, rate on a scale generates text: summaries, commands, rewrites compares vectors: nearest neighbours
output calibrated probabilities, labels, ranks, exit codes prose similarity scores
"angriest customer first" yes, that is a rubric you get a paragraph about anger no, anger is not a similarity
per line ~240 ms, ~$0.000013 (measured below) a round trip plus generated tokens, 19.3-257.9x the cost fast after indexing; needs an index
composes with &&, sort, head yes, by design awkwardly no
failure mode pass-through + exit 5; gate fails closed hallucinated text silent misses

The calibrated judgment layer for the shell. Generation is lx's job; similarity is grepai's. We judge, rank, classify, gate and route, and refuse to do anything else (CONTRIBUTING.md has the Hard Rule).

Numbers

All measured with the installed commands against a real endpoint, uncached; scripts and raw results are in bench/. Details and method notes: docs/benchmarks.md.

latency

lines judged per second

cost per 1000 decisions

accuracy

  • Latency: 238 ms median per decision end to end (p95 378 ms), measured through vercel; asking 16 questions about the same line costs about the same time as one.
  • Cost: 302 input tokens and $0.000013 per decision, $0.0127 per 1,000.
  • vs chat models: the same yes/no decision costs 19.3x to 257.9x more at list price (claude-fable-5.1, gpt-6-astra, claude-opus-5, …). Their latency is not measured here; see the method notes.
  • Accuracy, from a one-line description, with no tuning: jgrep finds SMS spam with F1 0.91 over all 5,574 messages, where a 17-term keyword regex scores 0.72 on the same text; jsort ranks review sentiment with AUC 1.00 over all 3,000 sentences; jtag labels AG News four ways with 87% accuracy over all 7,600 articles.

Every number above comes from a result file in bench/results/; the full tables and the method notes are in docs/benchmarks.md.

Two structural facts behind the numbers: Jev answers every question about one state in one call (so jgrep -e a -e b -e c and juniq's 50 pair questions cost tokens, not time), and Jev charges for input tokens only (about 300 per line; a million lines is about $13 at list price).

Common options

Every tool:

-p P             probability needed for a positive verdict (default 0.5)
--json           one JSON object per record, with scores and probabilities
--dry-run        print the backend, the exact questions and a redacted input sample; send nothing
--model ID       pin a model (default: the backend's latest Jev alias)
--api NAME       typesafe | openrouter | vercel | gateway
-j N             requests in flight (default 20, or $JEV_CONCURRENCY)
--timeout SEC    per request once it starts, retries and rate-limit waits included, but not
                 time spent queued behind -j (default 15, or $JEV_TIMEOUT)
--budget USD     stop at this spend (default 1.00, or $JEV_BUDGET; 0 = no limit)
--max-chars N    judge only the first N characters of a record (default 8000)
--no-cache       do not read or write ~/.cache/jev/answers.sqlite
--strict         fail closed: stop with exit 4 on the first API error
--color MODE     auto | always | never
--stats          calls, cached, tokens, dollars, p50 latency on stderr (default: when stderr is a TTY)
-q               no end-of-run notes on stderr; errors are still reported
-v               one line on stderr per decision, as it is made

jgrep keeps grep's -q instead: print nothing, stop at the first match, let the exit status be the answer. It also keeps grep's -v for invert-match, and spells the shared one --verbose.

Exit codes: 0 ok · 1 nothing matched · 2 usage or file error · 3 no or bad key · 4 API error or budget spent · 5 partial (some lines could not be judged and passed through).

CSV and JSONL in, CSV and JSONL out

A spreadsheet export is not a list of lines. Read as lines, its header row costs a real decision, comes back with a verdict on it, and lands wherever that verdict put it — so what you get out is no longer a file anything will open. --csv and --jsonl make the record the unit instead:

$ jsort "how urgent this ticket is" --csv examples/tickets.csv
id,customer,plan,message
1004,hooli,enterprise,"I was charged twice for one order, and nobody answers"
1006,acme,enterprise,Login fails with error 500 since this morning
1008,wayne,free,"Your support is a joke, three days without a reply"
1001,acme,enterprise,The app crashes when I open settings
…

A whole record goes to Jev as an object, so one description can weigh several columns at once ("a paying customer who is blocked" reads the plan and the message together); --field NAME narrows it to one value, and takes a dotted path in JSON. jtag adds its verdict as a real column or JSON key, jroute writes bug.csv and ask.csv each with their header, jmatch joins two exports whose columns are named differently (--field / --field-b).

Full behaviour, per tool, in docs/structured.md.

How it works

pipeline

  • Batching. All questions about one state travel in one request.
  • Concurrency. A semaphore bounds requests in flight (-j, default 20); the reader never runs far ahead of judging, so tail -f never buffers and a 10 GB file never loads.
  • Caching. Identical (model, state, question) is judged once per run, and once across runs in ~/.cache/jev/answers.sqlite (shared by all tools in a pipeline). Answers are stored under the model version that produced them, never under a moving alias like jev-latest, so a rerun spends one call to see what the alias means today and reads the rest from disk. Pin --model and a rerun costs nothing at all.
  • Fail-open. An API error means a line passes through unjudged, the run continues, the exit status becomes 5, and the error is reported once a minute. --strict fails closed. jgate fails closed by default, because a gate that opens during an outage is not a gate.
  • Rate limits. A 429 brakes every request in the client, and for the next two minutes requests go out one at a time instead of as a burst. Nothing fans out. A request waits the brake out when --timeout allows; when the backend asks for longer than that, the request says so immediately rather than sleeping out its whole deadline and reporting nothing useful.
  • Budget. --budget (default $1) stops a run before it costs more than you meant. The cost of a call is reserved when it starts, not billed when it ends, so twenty requests in flight cannot all clear the last dollar between them.
  • Two wire formats, one library. jevcore speaks TypeSafe's System One API (also OpenRouter and gateways) and the Vercel AI Gateway's evaluation modality; every tool sees typed answers only.

jevcore is importable on its own:

from jevcore import Jev, Noul, Choice, Score, resolve

async with Jev(resolve()) as jev:
    a = await jev.ask("checkout failed three times, I am done", {
        "angry":   Noul("The writer is frustrated."),
        "route":   Choice("Route this ticket.", {"billing": "payments", "bug": "defects"}),
        "urgency": Score("How urgent is this?", ["low", "medium", "high"]),
    })
    a["angry"].probability, a["route"].choice, a["urgency"].normalized   # 0.97, "bug", 0.93

More in docs/architecture.md.

Things to know

  • These are a model's judgments. Check a sample before you rely on a filter. Borderline lines get probabilities in the middle; that is what -p is for.
  • Jev answers the description you wrote, not the one you meant. TypeSafe documents weak spots: counting, comparing numbers or dates, double negatives, long inputs full of irrelevant detail.
  • Near-deterministic, not exactly: uncached probabilities can move by a few hundredths. Pin --model jev-1.13.0 and keep the cache for exact reruns.
  • Text in the input can try to steer the answer. Not a security boundary.
  • Free-tier gateways throttle hard; set JEV_TIMEOUT=600 and -j 2 and let the tools wait.

More in docs/faq.md.

Development

git clone https://github.com/adiel-hub/Jtools && cd Jtools
uv sync --all-groups
uv run pytest -q -m "not live"         # the offline suite against MockJev, no key needed
uv run pytest -q -m live               # one real call per tool, with any key set
uv run ruff check src tests && uv run ruff format --check src tests && uv run mypy src

The mock speaks both wire formats and runs as a real localhost endpoint for subprocess tests, so streaming, broken pipes, dead endpoints, ordering under latency and bounded concurrency are all covered without a key. CONTRIBUTING.md has the layout, the rules and the list of tools we deliberately do not build.

Acknowledgements

The idea of a shell tool whose pattern is a description, and the key/gateway conventions (~/.config/jev, JEV_API, JEV_GATEWAY_URL), come from Khaled Eltokhy's jgrep; our jgrep is a fresh implementation on the shared jevcore client and interoperates with the same config and cache. Jev is made by TypeSafe AI.

MIT license.

Metadata

Release files for jev-tools 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jev-tools 0.1.1
File Size Uploaded
jev_tools-0.1.1.tar.gz 846.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jev-tools 0.1.1
File Interpreter ABI Platform
jev_tools-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 951.8 kB

Release files / jev_tools-0.1.1.tar.gz

Download URL jev_tools-0.1.1.tar.gz
Size 846.1 kB
Tags Source
SHA-256 checksum
How to use checksums
7939bea224477f0e2295444a14785fdc301348cef7608990aff678f94938363e
BLAKE2b-256 checksum
How to use checksums
ecf1819f8ea8449be6831072df866beab632f8deb01aaeaf0dee2ef0f46c73bd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.17 {"installer":{"name":"uv","version":"0.12.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / jev_tools-0.1.1-py3-none-any.whl

Download URL jev_tools-0.1.1-py3-none-any.whl
Size 105.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
77822e7497947d1408dc7f74f6c39ec8199cac12523b0a992ee5f9688010f7d4
BLAKE2b-256 checksum
How to use checksums
e9914fe83bca7462673eb7a070a5b5b21d5e36bade5d1d8dee901100ca070f8c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.17 {"installer":{"name":"uv","version":"0.12.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page