SemBaker
Compile-then-execute semantic operators — an external optimizer for Palimpzest, LOTUS, Nirvana, and DocETL.
Semantic-operator systems let you write data operations in natural language —
sem_filter("clearly positive review"), sem_map("extract the patient's
age"), sem_join("opposite sentiment"). The usual way to execute these is
interpretation: for every row (or candidate pair, or document) the system
issues an LLM call to evaluate the predicate. Expressive, but it puts an
expensive LLM call inside the data loop — high latency, high cost, poor
scaling.
SemBaker does the alternative: compilation. It calls the LLM once to translate a semantic operator into a deterministic Python function, then runs that function locally over the whole dataset — no per-item LLM call at execution.
interpretation: N items -> N LLM calls (cost/latency grow with N)
compilation: N items -> 1 compile call + N local function calls (~0)
The compiled function is a reusable, cacheable artifact; the LLM cost is a fixed one-shot "compile" charge that amortizes across the input.
SemBaker implements the approach introduced in our vision paper From Interpretation to Compilation: Compilation-Based Execution of Semantic Operators (Dong & Wang, 2026) — see Citation.
Install
pip install sembaker # core only (compile / cache / decide)
pip install sembaker[pz] # + Palimpzest adapter deps
pip install sembaker[lotus] # + LOTUS (lotus-ai)
pip install sembaker[nirvana] # + Nirvana (nirvana-ai)
pip install sembaker[docetl] # + DocETL
pip install sembaker[all] # everything
Set OPENAI_API_KEY in your environment (or a .env you load yourself). The
default compile model is gpt-5-mini; override with CX_COMPILE_MODEL.
Usage: keep your engine's native API
You keep writing each engine's native pipeline; SemBaker is one extra call that walks the plan, decides per operator whether to compile-then-execute or stay native, and swaps in the compiled path — without modifying the engine's source.
Note: the cost model only picks the compiled path once the input is large enough to amortize the one-shot compile (roughly a dozen rows; see "Compile-vs-native decision" below). On toy-sized data everything routes native — the runnable scripts in examples/ use 20 rows so the compiled path actually engages.
Palimpzest
import palimpzest as pz
import sembaker.backends.pz as cxpz
ds = (pz.MemoryDataset(id="reviews", vals=df)
.sem_filter("the review is clearly positive"))
report = cxpz.optimize(ds) # walk plan, decide per op, gate rules
out = ds.run(config) # PZ executes; compiled ops run locally
LOTUS
import lotus
from lotus.ast import LazyFrame
import sembaker.backends.lotus as cxlotus
lf = LazyFrame(df).sem_filter("The {reviewText} is clearly positive.")
cxlotus.optimize(lf) # SemFilterNode -> CXFilterNode, in place
out = lf.execute(df)
Nirvana
import nirvana as nv
import sembaker.backends.nirvana as cxnv
ndf = nv.DataFrame(df)
ndf.semantic_filter("the review is clearly positive", input_columns=["reviewText"])
cxnv.optimize(ndf) # inject compiled fn into the native UDF slot
out, cost, secs = ndf.execute() # automatic LLM fallback if the fn raises
DocETL
import docetl
import sembaker.backends.docetl as cxdoc
cxdoc.apply() # wrap Frame.map / Frame.filter
f = docetl.from_list(docs)
f = f.filter(prompt="Keep clearly positive reviews. {{ input.reviewText }}")
out = f.collect() # rewritten to native code_filter, 0 LLM/doc
| Backend | How it's wired (no engine-source edits) | Module |
|---|---|---|
| Palimpzest | Compiled* physical operators + implementation rules |
sembaker.backends.pz, sembaker.backends.pz_ops |
| LOTUS | LazyFrame node rewrite (Sem*Node → CX*Node) |
sembaker.backends.lotus |
| Nirvana | native UDF-slot injection | sembaker.backends.nirvana |
| DocETL | Frame.map/filter → native code_map/code_filter |
sembaker.backends.docetl |
Your own engine: the compile core is public API; an adapter is typically ~100 lines. See docs/writing_a_backend.md.
Operator support
Every pipeline still runs in full. sembaker's decision is strictly per-operator: an operator it can't (or shouldn't) compile is simply left on the engine's native path — it is reported in the rewrite report, never rewritten, never broken. What CAN be compiled today is filter / map / join:
| Backend | filter | map | join | everything else |
|---|---|---|---|---|
| Palimpzest | ✅ | ✅ | ✅ (semantic condition; on= equi-joins already run natively without LLM) |
native (e.g. sem_agg) |
| LOTUS | ✅ | ✅ | ✅ inner joins only | native (sem_agg, sem_topk, sem_extract, ...) |
| Nirvana | ✅ | ✅ | ✅ | native (rank, reduce) |
| DocETL | ✅ | ✅ | — native (DocETL's equijoin has no code slot to inject into) |
native (reduce, resolve, ...) |
Also left native by design: operators where the user already supplied their own UDF (never touched), and predicates the codifiability judge rejects (e.g. tone/sarcasm/nuance judgments that keyword-or-regex logic can't faithfully capture).
The compile pipeline & knobs
Every backend routes its compiles through one entry point
(sembaker.optimizer.compile_op.compile_operator), so these environment switches
apply uniformly across all backends:
| Env switch | Values | Meaning |
|---|---|---|
CX_COMPILE |
e2e (default) · ir1 · ir2 |
compile method (see IR section) |
CX_CACHE |
1 (default) · 0 |
reuse/store compiled artifacts |
CX_REFINE |
0 (default) · 1 |
rewrite the predicate into a concrete, column-grounded one first |
CX_VALIDATE |
0 (default) · 1 |
pre-cache validation gate: score draws on LLM-labeled samples, keep the best |
CX_IR_FALLBACK |
1 |
if the IR can't express an operator, fall back to e2e |
CX_WARM |
0 (default) · 1 |
warm phase: pre-compile operators in parallel before execution |
CX_COMPILE_MODEL |
model id | LLM used for the one-shot compile (default gpt-5-mini) |
CX_JUDGE_MODEL |
model id | LLM used by the codifiability judge (defaults to CX_COMPILE_MODEL) |
- refine (
sembaker.core.refine) — one LLM call that turns a fuzzy predicate into a concrete recipe (e.g. discovers a structuredscoreSentimentcolumn to compare instead of guessing sentiment from text). - validate (
sembaker.core.validate) — labels a few sample items once (LLM, in parallel) and scores each compile draw against them, keeping the best. - cache (
sembaker.core.cache) — persistent, keyed on(method, op, canonicalized predicate, columns, model). The key excludes the backend: an artifact compiled for one engine is reused by another.
IR-decoupled compilation (CX_COMPILE=ir1|ir2)
Free-form code generation is high-variance. The IR path decouples finding the logic from writing the code: one LLM call emits a small structured IR, then a deterministic transpiler (no LLM) renders it to Python — so temperature-1 variance is confined to a small structured space, and the code step is byte-reproducible.
ir1— a flat "feature + decision" DSL (sembaker.core.ir).ir2— a nestable expression tree (AST) produced via grammar prompting; richer (nesting, comparisons,to_int,first_nonempty,ifelse) yet still one LLM call (sembaker.core.ir2). Design follows TRANX (Yin & Neubig, EMNLP 2018) with grammar prompting (Wang et al., 2023).
Compile-vs-native decision
sembaker.optimizer.decide() routes each operator: a codifiability judge
(heuristic or LLM) asks whether the predicate is expressible as deterministic
logic over the visible columns, and a cost model amortizes the one-shot
compile cost against the per-item native cost at the operator's estimated
cardinality. Force the compiled path with each adapter's force flag / the
FORCE_CX=1 convention where supported.
Tested versions
Developed and tested against: palimpzest 1.5.x · lotus-ai 1.1.x ·
nirvana-ai 1.3.x · docetl 0.3.x · Python 3.10+.
Citation
If you use SemBaker in your research, please cite:
@misc{dong2026compilation,
title = {From Interpretation to Compilation: Compilation-Based
Execution of Semantic Operators},
author = {Dong, Wenkai and Wang, Yifan},
year = {2026},
eprint = {2607.13407},
archivePrefix = {arXiv},
primaryClass = {cs.DB},
url = {https://arxiv.org/abs/2607.13407},
}
License
SemBaker is dual-licensed:
- AGPL-3.0 (LICENSE) — free for research, evaluation, and any use that complies with the AGPL's copyleft, including its network-service provision (running a modified version as a service requires releasing your corresponding source under the AGPL).
- Commercial license — for closed-source or proprietary use that cannot meet AGPL obligations. Contact dongw@hawaii.edu or yifanw@hawaii.edu.
Patent applications covering techniques in this software have been filed; see NOTICE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sembaker-0.1.0.tar.gz.
File metadata
- Download URL: sembaker-0.1.0.tar.gz
- Upload date:
- Size: 73.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
04322c7a4fa7b07fa6e5b86f0b73714572f41ddec3d98eba79f4d6593e126d4a
|
|
| MD5 |
eb0fbf48d8d7c463acf0966bba0a3034
|
|
| BLAKE2b-256 |
cfe56481dccee94c51a748d6d6b2356fe928015c1e569d57632240ef99afb751
|
File details
Details for the file sembaker-0.1.0-py3-none-any.whl.
File metadata
- Download URL: sembaker-0.1.0-py3-none-any.whl
- Upload date:
- Size: 87.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fc724a2210e7a02aaa64717247ac15b74e4c18a63fc804efba72659f1f863d6b
|
|
| MD5 |
b8f53efa92419fc1f0ad14d994420b93
|
|
| BLAKE2b-256 |
db7c6c2465d17adc15a87385e89d3a48265abe8a7527977e0309fd5ccd196f95
|