Skip to main content

Jev2SemOpt — Typed decisions over tables

Filter tickets. Assign queues. Rank records. Match pairs.

A support pipeline needs to keep refund requests, assign a fixed queue, and count the resulting tickets. Jev2SemOpt makes those dataset operations explicit: select the fields the model may see, bind a typed question once, evaluate each record, and apply the result with ordinary Pandas operations.

Jev2SemOpt is an independent, early-stage Python library inspired by LOTUS semantic operators, using LLM2Jev for Jev-style decisions. The Jev interface concepts and Choice, Score, Noul vocabulary originate with TypeSafe. Use a local LLM2Jev runtime or the official TypeSafe HTTP API through JevBackend. The HTTP adapter is contract-tested; hosted Jev accuracy has not been measured.

Install and run

Python 3.8+. The core requires Pandas and LLM2Jev; model providers are opt-in.

pip install jev2semopt
# Official hosted Jev (requires a TypeSafe API key):
pip install 'jev2semopt[jev]'
# Local providers:
pip install 'jev2semopt[transformers]'
# or: pip install 'jev2semopt[ollama]'

For a model-free example or development:

git clone https://github.com/Qingbolan/Jev2SemOpt.git
cd Jev2SemOpt
uv sync --group dev
uv run python examples/refund_pipeline.py
uv run python examples/decorated_decision.py

These examples use synthetic scores and require no model or API key. Production and development dependencies resolve from PyPI. See release verification.

A refund pipeline

Configure execution explicitly. The application owns the runtime and closes it; Jev2SemOpt only borrows it. The following requires compatible local model weights and a model-appropriate LLM2Jev encoder/label configuration; it is an API example, not a validated inference configuration.

import pandas as pd
from llm2jev import LLM2Jev, TransformersRuntime
from jev2semopt import LLM2JevBackend, SemEngine, Choice

tickets = pd.DataFrame({
    "id": [101, 102],
    "text": ["Please refund my order.", "Where is my parcel?"],
})

with TransformersRuntime("/path/to/local/model", device="cpu") as runtime:
    engine = SemEngine(LLM2JevBackend(
        LLM2Jev(runtime=runtime, model_identity=runtime.identity),
        model=runtime.identity.name,
    ))
    refunds = engine.sem_filter(
        tickets, "Does text explicitly request a refund?",
        columns=["text"], threshold=0.7, probability_column="refund_support",
    )
    routed = engine.sem_map(
        refunds,
        Choice(instructions="Which team should handle text?", criteria={
            "billing": "Payments, invoices, or refunds",
            "delivery": "Shipping, tracking, or missing parcels",
            "other": "Requests outside billing and delivery",
        }),
        columns=["text"], output="queue",
    )
    counts = routed.groupby("queue").size()

Only text enters the decision state. IDs remain available in the output. Column names in instructions refer to keys in that state; there is no {column} string interpolation. Thresholds require validation on your own labeled workload.

Operators and boundaries

Official hosted Jev is also supported through JevBackend, using a caller-owned HTTP client and TypeSafe API key. See the official API integration for setup and verification limits. An OpenRouter key does not authenticate TypeSafe.

Operator Jev decision Output and scope
sem_filter(frame, instructions, ...) Noul Rows with support ≥ threshold, original order and index
sem_map(frame, Choice(...), output=...) Choice All rows plus a chosen label from supplied alternatives
sem_score(frame, Score(...), ...) Score All rows plus expected ordinal rubric level
sem_topk(frame, Score(...), k=...) Score Largest expected levels; stable source-order ties
sem_join(left, right, instructions, ...) Noul Inner join over candidate pairs; namespaced source fields and positions

LOTUS provides a broader semantic operator model, including generated projections, extraction, aggregation, and comparator-based ranking. Jev2SemOpt deliberately restricts mappings to finite alternatives and ranks by an explicit ordinal rubric. It is not a drop-in LOTUS replacement. Free-text extraction, summaries, vector search, learned cascades, SQL planning, asynchronous execution, and distributed execution are not implemented. Count/group/sum the typed outputs with Pandas; there is no misleading sem_agg alias for a generative summary.

from jev2semopt import Score

ranked = engine.sem_topk(
    tickets,
    Score(instructions="How urgent is text?", criteria=[
        "Routine request", "Time-sensitive issue", "Immediate safety or service emergency",
    ]),
    columns=["text"], k=10, output="urgency",
)

matches = engine.sem_join(
    tickets, policies,
    "Does left.text describe a case covered by right.policy?",
    left_columns=["text"], right_columns=["policy"],
    candidates=[(0, 1), (1, 0)],  # row POSITIONS, not index labels
    threshold=0.8,
)

The join output includes left.<column>, right.<column>, _left_position, _right_position, and _probability. Without candidates it evaluates every pair, subject to a default 100,000-pair limit. A shortlist can reduce work but can also exclude true matches; the caller owns its recall. See API contracts.

Python-native integration

The Pandas accessor is opt-in and uses Pandas' registration decorator. It forwards to the same engine implementation and does not install global model settings:

import jev2semopt.pandas

refunds = tickets.jev.sem_filter(
    engine, "Does text request a refund?", columns=["text"], threshold=0.7,
)

Use @decision when application code already builds the state for one decision:

from jev2semopt import Noul, decision

@decision(engine, question=Noul(instructions="Does text explicitly request a refund?"))
def refund_requested(text):
    return {"text": text}

answer = refund_requested("Please refund my order.")
print(answer.noul)

The decorator binds once, preserves function metadata with functools.wraps, and evaluates new state on each call. It does not cache results across calls. Decorated functions must be synchronous JSON-state builders and now return typed answers.

Measured results

On 50 balanced SciFact records, the same GPT-4o-mini scored 31/50 with LOTUS and 32/50 with Jev-style / LLM2Jev. The one-record difference does not establish an accuracy improvement: the paired 95% interval spans −8 to +12 percentage points. Jev-style improves precision but lowers recall and F1 in this run.

Small-sample accuracy, precision, recall and F1 for LOTUS and Jev-style

On 400 local Qwen2.5-0.5B candidate pairs, Jev-style filtering was 1.97× faster, but both filters had roughly 4% precision. Three-level Jev-style scoring was 2.70× slower than LOTUS scoring and reduced ranking quality. BM25 had the highest nDCG.

Local ranking quality and operator time, including the BM25 baseline

These are official SciFact dataset subsets with adapted protocols, not full LOTUS paper reproduction or measurements of TypeSafe's hosted Jev model. Unjudged documents count as negatives under the qrel convention. Prompts differ between operators; these experiments do not isolate probability assembly as the cause. Sampling, confusion matrices, costs, raw observations, and reproduction.

Execution cost and score meaning

SemEngine(backend, deduplicate=True) reuses identical selected JSON state within one operation. This is opt-in and assumes the backend is deterministic and has no per-call side effects. There is no global cache and no stale reuse across operations. Question binding avoids recompilation; it is not a KV cache or a batched model call.

For N rows, filtering requires N one-candidate evaluations; mapping with C choices requires N × C binary candidates; scoring with R rubric levels requires N × R. An exhaustive join needs L × R pair evaluations. Duplicate reuse reduces these counts to unique selected states. Calls are sequential. Join candidates and results are materialized in memory; this version targets bounded in-memory tables.

No end-to-end speedup, calibrated correctness, or equivalence to LOTUS quality is claimed. Noul is label-conditioned support; Score is an expected equally spaced rubric index. These are decision signals, not probabilities that the answer is right. LLM2Jev's Transformers adapter reads next-token logits; its Ollama adapter requests one token to obtain exact binary logprobs. Jev2SemOpt never parses generated prose. See evaluation protocol.

Architecture and development

DataFrame API / @decision
          ↓
SemEngine — positional relational semantics
          ↓
EvaluationSession — one rule, detached state, operation-local reuse
          ↓
DecisionBackend.bind → BoundDecision.evaluate → typed Answer
          ↓
├─ LLM2JevBackend → public LLM2Jev API → caller-owned runtime
└─ JevBackend → TypeSafe HTTP API → caller-owned HTTP client

The engine depends on a backend protocol, not Ollama or Transformers. Models, prompts, device configuration, binary labels, and runtime lifecycle stay in LLM2Jev for local execution; hosted execution uses the supplied HTTP client. There is no new resource lifecycle to duplicate in the table layer.

src/jev2semopt/
├── contracts.py       Backend and bound-decision protocols
├── execution.py       State isolation, answer checks, safe failure boundary
├── engine.py          Filter/map/score/top-k/join semantics
├── table.py           Dataframe schema and state projection
├── decorators.py      Function-to-decision binding
├── pandas.py          Opt-in accessor registration
└── adapters/          LLM2Jev service and official Jev HTTP integration

Run the checks in CONTRIBUTING.md. Architectural decisions, privacy constraints, and delivery status are documented in architecture, privacy, and implementation plan. Attribution is recorded in NOTICE. Licensed under MIT; dependency and benchmark attribution.

Local verification record: 50 deterministic tests passed on Python 3.8, 3.12, and 3.14. Real-model observations are recorded separately in the benchmark report.

Metadata

Release files for jev2semopt 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jev2semopt 0.1.0
File Size Uploaded
jev2semopt-0.1.0.tar.gz 674.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jev2semopt 0.1.0
File Interpreter ABI Platform
jev2semopt-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 693.6 kB

Release files / jev2semopt-0.1.0.tar.gz

Download URL jev2semopt-0.1.0.tar.gz
Size 674.7 kB
Tags Source
SHA-256 checksum
How to use checksums
bd84fa34061ba1d7e572320c701419ef7be16f482db5958f3b9651bda2600325
BLAKE2b-256 checksum
How to use checksums
9f16539ef405d920a5b9551cd22a636a642cd6feb8efa51f2c2d216a99529b36
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release files / jev2semopt-0.1.0-py3-none-any.whl

Download URL jev2semopt-0.1.0-py3-none-any.whl
Size 18.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2658bcdbcb88d57d910eec7d8ae5aef3b20aba31f645ebe869333a82193110b4
BLAKE2b-256 checksum
How to use checksums
b10b67efe951864ee44ddc7bb8c9f11882573f95d129f9a800b65a77d3fe952b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release history Release notifications | RSS feed

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page