Jev2SemOpt — Typed decisions over tables
Filter tickets. Assign queues. Rank records. Match pairs.
A support pipeline needs to keep refund requests, assign a fixed queue, and count the resulting tickets. Jev2SemOpt makes those dataset operations explicit: select the fields the model may see, bind a typed question once, evaluate each record, and apply the result with ordinary Pandas operations.
Jev2SemOpt is an independent, early-stage Python library inspired by
LOTUS semantic operators, using
LLM2Jev for Jev-style decisions.
The Jev interface concepts and Choice, Score, Noul vocabulary originate with
TypeSafe.
Use a local LLM2Jev runtime or the official TypeSafe HTTP API through JevBackend.
The HTTP adapter is contract-tested; hosted Jev accuracy has not been measured.
Install and run
Python 3.8+. The core requires Pandas and LLM2Jev; model providers are opt-in.
pip install jev2semopt
# Official hosted Jev (requires a TypeSafe API key):
pip install 'jev2semopt[jev]'
# Local providers:
pip install 'jev2semopt[transformers]'
# or: pip install 'jev2semopt[ollama]'
For a model-free example or development:
git clone https://github.com/Qingbolan/Jev2SemOpt.git
cd Jev2SemOpt
uv sync --group dev
uv run python examples/refund_pipeline.py
uv run python examples/decorated_decision.py
These examples use synthetic scores and require no model or API key. Production and development dependencies resolve from PyPI. See release verification.
A refund pipeline
Configure execution explicitly. The application owns the runtime and closes it; Jev2SemOpt only borrows it. The following requires compatible local model weights and a model-appropriate LLM2Jev encoder/label configuration; it is an API example, not a validated inference configuration.
import pandas as pd
from llm2jev import LLM2Jev, TransformersRuntime
from jev2semopt import LLM2JevBackend, SemEngine, Choice
tickets = pd.DataFrame({
"id": [101, 102],
"text": ["Please refund my order.", "Where is my parcel?"],
})
with TransformersRuntime("/path/to/local/model", device="cpu") as runtime:
engine = SemEngine(LLM2JevBackend(
LLM2Jev(runtime=runtime, model_identity=runtime.identity),
model=runtime.identity.name,
))
refunds = engine.sem_filter(
tickets, "Does text explicitly request a refund?",
columns=["text"], threshold=0.7, probability_column="refund_support",
)
routed = engine.sem_map(
refunds,
Choice(instructions="Which team should handle text?", criteria={
"billing": "Payments, invoices, or refunds",
"delivery": "Shipping, tracking, or missing parcels",
"other": "Requests outside billing and delivery",
}),
columns=["text"], output="queue",
)
counts = routed.groupby("queue").size()
Only text enters the decision state. IDs remain available in the output. Column
names in instructions refer to keys in that state; there is no {column} string
interpolation. Thresholds require validation on your own labeled workload.
Operators and boundaries
Official hosted Jev is also supported through JevBackend, using a caller-owned
HTTP client and TypeSafe API key. See the official API integration
for setup and verification limits. An OpenRouter key does not authenticate TypeSafe.
| Operator | Jev decision | Output and scope |
|---|---|---|
sem_filter(frame, instructions, ...) |
Noul |
Rows with support ≥ threshold, original order and index |
sem_map(frame, Choice(...), output=...) |
Choice |
All rows plus a chosen label from supplied alternatives |
sem_score(frame, Score(...), ...) |
Score |
All rows plus expected ordinal rubric level |
sem_topk(frame, Score(...), k=...) |
Score |
Largest expected levels; stable source-order ties |
sem_join(left, right, instructions, ...) |
Noul |
Inner join over candidate pairs; namespaced source fields and positions |
LOTUS provides a broader semantic operator model, including generated projections,
extraction, aggregation, and comparator-based ranking. Jev2SemOpt deliberately
restricts mappings to finite alternatives and ranks by an explicit ordinal rubric.
It is not a drop-in LOTUS replacement. Free-text extraction, summaries, vector
search, learned cascades, SQL planning, asynchronous execution, and distributed
execution are not implemented. Count/group/sum the typed outputs with Pandas;
there is no misleading sem_agg alias for a generative summary.
from jev2semopt import Score
ranked = engine.sem_topk(
tickets,
Score(instructions="How urgent is text?", criteria=[
"Routine request", "Time-sensitive issue", "Immediate safety or service emergency",
]),
columns=["text"], k=10, output="urgency",
)
matches = engine.sem_join(
tickets, policies,
"Does left.text describe a case covered by right.policy?",
left_columns=["text"], right_columns=["policy"],
candidates=[(0, 1), (1, 0)], # row POSITIONS, not index labels
threshold=0.8,
)
The join output includes left.<column>, right.<column>, _left_position,
_right_position, and _probability. Without candidates it evaluates every pair,
subject to a default 100,000-pair limit. A shortlist can reduce work but can also
exclude true matches; the caller owns its recall. See API contracts.
Python-native integration
The Pandas accessor is opt-in and uses Pandas' registration decorator. It forwards to the same engine implementation and does not install global model settings:
import jev2semopt.pandas
refunds = tickets.jev.sem_filter(
engine, "Does text request a refund?", columns=["text"], threshold=0.7,
)
Use @decision when application code already builds the state for one decision:
from jev2semopt import Noul, decision
@decision(engine, question=Noul(instructions="Does text explicitly request a refund?"))
def refund_requested(text):
return {"text": text}
answer = refund_requested("Please refund my order.")
print(answer.noul)
The decorator binds once, preserves function metadata with functools.wraps, and
evaluates new state on each call. It does not cache results across calls. Decorated
functions must be synchronous JSON-state builders and now return typed answers.
Measured results
On 50 balanced SciFact records, the same GPT-4o-mini scored 31/50 with LOTUS and 32/50 with Jev-style / LLM2Jev. The one-record difference does not establish an accuracy improvement: the paired 95% interval spans −8 to +12 percentage points. Jev-style improves precision but lowers recall and F1 in this run.
On 400 local Qwen2.5-0.5B candidate pairs, Jev-style filtering was 1.97× faster, but both filters had roughly 4% precision. Three-level Jev-style scoring was 2.70× slower than LOTUS scoring and reduced ranking quality. BM25 had the highest nDCG.
These are official SciFact dataset subsets with adapted protocols, not full LOTUS paper reproduction or measurements of TypeSafe's hosted Jev model. Unjudged documents count as negatives under the qrel convention. Prompts differ between operators; these experiments do not isolate probability assembly as the cause. Sampling, confusion matrices, costs, raw observations, and reproduction.
Execution cost and score meaning
SemEngine(backend, deduplicate=True) reuses identical selected JSON state within
one operation. This is opt-in and assumes the backend is deterministic and has no
per-call side effects. There is no global cache and no stale reuse across operations.
Question binding avoids recompilation; it is not a KV cache or a batched model call.
For N rows, filtering requires N one-candidate evaluations; mapping with C
choices requires N × C binary candidates; scoring with R rubric levels requires
N × R. An exhaustive join needs L × R pair evaluations. Duplicate reuse reduces
these counts to unique selected states. Calls are sequential. Join candidates and
results are materialized in memory; this version targets bounded in-memory tables.
No end-to-end speedup, calibrated correctness, or equivalence to LOTUS quality is
claimed. Noul is label-conditioned support; Score is an expected equally spaced
rubric index. These are decision signals, not probabilities that the answer is right.
LLM2Jev's Transformers adapter reads next-token logits; its Ollama adapter requests
one token to obtain exact binary logprobs. Jev2SemOpt never parses generated prose.
See evaluation protocol.
Architecture and development
DataFrame API / @decision
↓
SemEngine — positional relational semantics
↓
EvaluationSession — one rule, detached state, operation-local reuse
↓
DecisionBackend.bind → BoundDecision.evaluate → typed Answer
↓
├─ LLM2JevBackend → public LLM2Jev API → caller-owned runtime
└─ JevBackend → TypeSafe HTTP API → caller-owned HTTP client
The engine depends on a backend protocol, not Ollama or Transformers. Models, prompts, device configuration, binary labels, and runtime lifecycle stay in LLM2Jev for local execution; hosted execution uses the supplied HTTP client. There is no new resource lifecycle to duplicate in the table layer.
src/jev2semopt/
├── contracts.py Backend and bound-decision protocols
├── execution.py State isolation, answer checks, safe failure boundary
├── engine.py Filter/map/score/top-k/join semantics
├── table.py Dataframe schema and state projection
├── decorators.py Function-to-decision binding
├── pandas.py Opt-in accessor registration
└── adapters/ LLM2Jev service and official Jev HTTP integration
Run the checks in CONTRIBUTING.md. Architectural decisions, privacy constraints, and delivery status are documented in architecture, privacy, and implementation plan. Attribution is recorded in NOTICE. Licensed under MIT; dependency and benchmark attribution.
Local verification record: 50 deterministic tests passed on Python 3.8, 3.12, and 3.14. Real-model observations are recorded separately in the benchmark report.
Metadata
Release files for jev2semopt 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| jev2semopt-0.1.0.tar.gz | 674.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| jev2semopt-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 693.6 kB
Release files / jev2semopt-0.1.0.tar.gz
| Download URL | jev2semopt-0.1.0.tar.gz |
|---|---|
| Size | 674.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bd84fa34061ba1d7e572320c701419ef7be16f482db5958f3b9651bda2600325
|
|
BLAKE2b-256 checksum How to use checksums |
9f16539ef405d920a5b9551cd22a636a642cd6feb8efa51f2c2d216a99529b36
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.7
|
Release files / jev2semopt-0.1.0-py3-none-any.whl
| Download URL | jev2semopt-0.1.0-py3-none-any.whl |
|---|---|
| Size | 18.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2658bcdbcb88d57d910eec7d8ae5aef3b20aba31f645ebe869333a82193110b4
|
|
BLAKE2b-256 checksum How to use checksums |
b10b67efe951864ee44ddc7bb8c9f11882573f95d129f9a800b65a77d3fe952b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.7
|