Skip to main content

jev-pandas

Explore a pandas dataframe using natural-language judgments from Jev (TypeSafe System One) compatible endpoints. Find incidents by meaning, classify records, or score them against a rubric — from a notebook or a script. No generative chat model needed.

flowchart LR
    DF["pandas DataFrame"] --> JF["JevFrame(df, client)"]
    CL["JevClient(base_url, api_key, model)"] --> JF
    JF --> EV["evaluate(condition)"]
    JF --> FI["filter(condition, threshold)"]
    JF --> CLS["classify(question, choices)"]
    JF --> SC["score(question, levels)"]
    JF --> ASK["ask({name: noul | choice | score})"]
    EV --> RES["JevResult"]
    FI --> RES
    CLS --> RES
    SC --> RES
    ASK --> RES
    RES --> OUT["original rows + judgment columns, metadata"]

Install

Requires Python 3.10+ and uv. The PyPI distribution is jev-pandas; the import name is jevpandas.

uv sync --extra dev
export TYPESAFE_BASE_URL=https://api.typesafe.ai/v1
export TYPESAFE_API_KEY=your-key
export TYPESAFE_MODEL=jev-latest

TYPESAFE_BASE_URL and TYPESAFE_MODEL default to the official Jev endpoint (https://api.typesafe.ai/v1, jev-latest). Any TypeSafe System One compatible endpoint works; pass base_url, api_key, and model to JevClient to override.

For notebook work (JupyterLab, ipykernel, and the optional tqdm_progress() bar), install the notebook extra as well: uv sync --extra dev --extra notebook.

Notebook

import pandas as pd
from jevpandas import JevClient, JevFrame, noul, choice, score, tqdm_progress

df = pd.read_parquet(
    "data/incidents.parquet"
)  # 1,000 synthetic rows; read_csv("data/incidents.csv") works too

with JevClient() as client:
    incidents = JevFrame(df, client)

    # All rows, including probabilities and explicit row errors.
    evaluated = incidents.evaluate(
        "Customers cannot access the service and the problem remains unresolved.",
        columns=["subject", "message"],
        workers=4,  # process rows concurrently, order preserved
        progress=tqdm_progress(),  # optional progress bar
    )
    matches = incidents.filter(
        "Customers cannot access the service and the problem remains unresolved.",
        columns=["subject", "message"],
        threshold=0.7,
        workers=4,
    )
    print(matches.groupby("service").size())

    classified = incidents.classify(
        "What is the main topic?",
        choices={
            "access": "Login or authentication",
            "billing": "Payments or refunds",
            "technical": "Other service problems",
            "other": "Anything else",
        },
        columns=["subject", "message"],
    )
    scored = incidents.score(
        "How severely does this currently disrupt customers?",
        levels=["No current disruption", "Partly impaired", "Unable to use the service"],
        columns=["subject", "message"],
    )

evaluate, classify, and score return a JevResult (a pandas.DataFrame subclass) with additional columns and a run summary in its HTML repr. Index order and duplicate index labels are preserved. Use name= to choose an output prefix and avoid collisions when adding multiple judgments. filter raises if any row fails; use evaluate to inspect partial results. Missing selected values are encoded as JSON null; entirely empty rows are reported as errors without making a model call. Score values range from zero to len(levels)-1.

Ask several questions in one pass

ask sends every question for a row in a single request, which is much cheaper than a separate pass per question:

result = incidents.ask(
    {
        "access": noul("Is this a current, unresolved customer access problem?"),
        "topic": choice(
            "What is the main topic?", {"access": "Login", "billing": "Payments", "other": "Other"}
        ),
        "severity": score("How severe is the current disruption?", ["None", "Impaired", "Outage"]),
    },
    columns=["subject", "message"],
    workers=4,
)

Each question produces {name}_{field} columns: access_probability, topic_label, severity_value, plus {name}_error, {name}_model, {name}_cached per question.

Concurrency and caching

Rows run sequentially by default (workers=1); pass workers=N to evaluate them concurrently with a thread pool. Results are always assembled in the original row order, and a failed row never affects the others. The client retries transient network errors, HTTP 429, and server errors; failures are never turned into negative predictions. Identical successful requests are cached in memory (up to 10,000 entries per client/session), guarded by a lock so parallel runs are safe. Changing the instruction, selected data, endpoint, or requested model invalidates the relevant cache entry. Threshold changes need no requests. Pin the model version, or call clear_cache() when an alias changes. No API keys or datasets are stored in a disk cache.

Notebook display

JevResult renders with a summary line (rows, elapsed time, cache hits, errors, model) above the table. Run metadata is also available as result.metadata and stored in result.attrs["jev"].

Notes

Model probabilities and confidence are backend-reported values, not verified accuracy guarantees. OpenJev and hosted Jev have different models and must be evaluated separately. This version does not translate arbitrary chat into code, generate explanations, run joins, or train on review labels. Samples and their expected labels in data/ are synthetic and are not evidence of accuracy on production data.

Validate

uv run pytest
uv run ruff check .
# Calls the configured server on 1,000 synthetic rows and writes ignored output/ files:
uv run python scripts/evaluate_sample.py

The live script exports all predictions, matching incidents, and a summary with a confusion matrix, timing, returned model IDs, and a check that repeating the run uses cached decisions. Expected labels are kept in data/incident_labels.csv and are never sent to the model.

Protocol references

The current local server advertises OpenJev 0.1 backed by DiffusionGemma 26B-A4B. Its /openapi.json defines the protocol used here; no assumptions about the upstream repository's current model are needed.

Release files for jev-pandas 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jev-pandas 0.1.1
File Size Uploaded
jev_pandas-0.1.1.tar.gz 62.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jev-pandas 0.1.1
File Interpreter ABI Platform
jev_pandas-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 76.1 kB

Release files / jev_pandas-0.1.1.tar.gz

Download URL jev_pandas-0.1.1.tar.gz
Size 62.8 kB
Tags Source
SHA-256 checksum
How to use checksums
83c2611c5c4e943ffbbb9891042ee125debe82624875c0410a616a4ed172d9cb
BLAKE2b-256 checksum
How to use checksums
0692feec6acf1e0022922255d8475b55112b210d6059c00cd4f6fc8540b584f9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / jev_pandas-0.1.1-py3-none-any.whl

Download URL jev_pandas-0.1.1-py3-none-any.whl
Size 13.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2d551ea1ed40011c9f28c2e0a30308633a30ce7c55d23ba2d82d1e1e2e1d39e4
BLAKE2b-256 checksum
How to use checksums
9cda445aff4f438f181a2b763dd7d89c85d2aa215de30869a3cd71115bcb4e52
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page