SIMIT-ICL
Vision-language models that imagine their own in-context demonstrations at test time.
Given an unlabeled query (image, question), SIMIT-ICL has the model synthesize a few similar
(image, question, answer) examples, verifies them, keeps only the useful ones, and answers the query
with them in context. No labels, no extra training, and one model does all the work (or a standard VLM
plus a text-to-image tool).
from simit import SIMIT
model = SIMIT.from_pretrained("ByteDance-Seed/BAGEL-7B-MoT")
demos = model.imagine(image, question) # imagined (image, question, answer) demos
answer = model.answer(image, question, demos) # SIMIT-ICL answer
greedy = model.greedy(image, question) # standard zero-shot answer, for comparison
How one query is processed (paper, Sec. 3 / App. G):
- Zero-shot pass: greedy answer + confidence
p0(geometric-mean token probability). - Adaptive budget (ABA):
K*(p0)demos, 0 for queries the model is already sure about. - Triplet synthesis: the model proposes
(description, question, answer)triplets similar to the query. - Realization: a router picks one of 17 skills. Natural images come from the model's own image head (BAGEL, Lance) or an image-generator tool (standard VLMs). Charts, diagrams, molecules, circuits, tables, flowcharts, etc. are written as code/specs by the model and rendered deterministically, with parse/render errors fed back for repair.
- Critic: the model scores each image against its description (0-100). Failures are revised and regenerated.
- Difficulty filter (DF): keep demos whose teacher-forced answer confidence lies in
[t_low, t_high]. - ICL answer with the kept demos in context.
Installation
pip install simit # BAGEL, Lance, any transformers VLM
pip install "simit[vllm]" # + the vLLM engine for standard VLMs (much faster than plain transformers)
# or from source
git clone https://github.com/monurcan/simit && cd simit && pip install -e .
Every skill dependency is a pip wheel: there is no Node.js, Mermaid CLI, cairo or manual browser setup. HTML and Mermaid skills render in headless Chromium through Playwright, with the following fallbacks:
- Playwright's own browser, if one is installed.
- A system
chromium/chrome. - Otherwise SIMIT downloads Playwright's headless Chromium once on first use (~100 MB, file-locked so parallel processes don't race). If Playwright's downloader fails, as it does behind some HPC networks, it fetches the same build with plain Python (proxy environment variables are honored).
Mermaid itself ships inside the package.
Requirements: Python ≥ 3.10, PyTorch ≥ 2.12 (BAGEL/Lance use torch.nn.attention.varlen, so no
flash-attn build is needed), one CUDA GPU (H100/A100 80GB for the 27B example; BAGEL-7B and Lance
fit on smaller cards).
Supported models
| model | how images are made | engine |
|---|---|---|
ByteDance-Seed/BAGEL-7B-MoT |
native image head + 17 skills, critic in thinking mode (as in the paper) | built-in MoT engine |
bytedance-research/Lance |
native image head; defaults to decomposed synthesis, natural images, no critic (its instruction following is weaker) | built-in MoT engine |
any transformers VLM, e.g. Qwen/Qwen3.8-27B |
image_generator= tool (e.g. black-forest-labs/FLUX.2-klein-4B) + 17 skills; omit the tool for skills only |
vLLM if installed, else transformers |
# a standard VLM + a text-to-image tool
model = SIMIT.from_pretrained("Qwen/Qwen3.8-27B", image_generator="black-forest-labs/FLUX.2-klein-4B")
# force an engine
model = SIMIT.from_pretrained("Qwen/Qwen3.8-27B", engine="transformers") # or "vllm"
# wrap a model you already loaded with the standard HF interface
from transformers import AutoProcessor, AutoModelForMultimodalLM
processor = AutoProcessor.from_pretrained("Qwen/Qwen3.8-27B")
hf_model = AutoModelForMultimodalLM.from_pretrained("Qwen/Qwen3.8-27B", device_map="auto")
model = SIMIT.from_pretrained(hf_model, processor=processor, image_generator="black-forest-labs/FLUX.2-klein-4B")
image_generator can be a diffusers model id, a loaded diffusers pipeline, or any callable
(prompt, width, height) -> PIL.Image. On one GPU it is loaded first and the VLM engine sizes its
memory around it. With several GPUs it goes on the last one.
Usage
from simit import SIMIT, save_demos
model = SIMIT.from_pretrained("ByteDance-Seed/BAGEL-7B-MoT")
# 1) fit ABA + DF on a small labeled validation set (synthesis runs once and is cached)
model.tune(val_set, metric="vqa_accuracy", max_runtime_ratio=25)
# 2) test time
for image, question in test_set:
demos = model.imagine(image, question)
improved = model.answer(image, question, demos)
Main methods:
| method | what it does |
|---|---|
imagine(image, question, k=None, return_details=False, time_limit=None, on_demo=None, on_progress=None) |
adaptive number of demos (ABA + DF); k=n asks for n verified demos (fewer if the attempt budget runs out first, e.g. when the critic keeps rejecting images); return_details=True returns an Imagination with the zero-shot answer, p0, the budget, every candidate and stats. time_limit (seconds) stops and returns the demos accepted so far; on_demo(demo) / on_progress(stage) stream progress, e.g. to a UI |
zero_shot(image, question) |
the greedy answer with its confidence p0; model.config.budget(zs.confidence) is the number of demos ABA will ask for. Reused by a following imagine |
answer(image, question, demos) |
answer with the demos in context (no demos: zero-shot) |
greedy(image, question) |
the standard zero-shot greedy answer (cached and shared with imagine) |
model(image, question) |
imagine then answer |
imagine_batch(queries) / answer_batch(queries, demos) |
many queries at once; their model calls are batched together on the GPU (much faster than a loop) |
tune(val_set, metric, n_trials=100, max_runtime_ratio=None, max_candidates=None, cache_dir=None, k_max_choices=None) |
hyperparameter search (below) |
add_skill(skill) / remove_skill(name) |
change the skill library |
val_set items are (image, question, answer_or_answers) tuples or dicts with
image/question/answer(s) (and optionally their own metric, for validation sets that mix benchmarks). metric is one of exact_match, vqa_accuracy, contains,
anls, relaxed_accuracy, multiple_choice, or any fn(prediction, references) -> float.
Images can be PIL images, file paths or raw bytes.
Saving demos for fine-tuning (SIMIT-FT)
from simit import save_demos, load_demos
save_demos(demos, "imagined/query_0") # PNGs + demos.jsonl (with LLaVA-style "conversations")
The JSONL has image, question, answer, description, skill, confidence,
verify_score and a conversations field, ready for LLaVA-format fine-tuning scripts. SIMIT-FT
itself (fine-tuning on the imagined test-set data) is out of scope for this package.
Hyperparameter tuning
tune follows App. G. For each validation query it runs synthesis once and caches the
zero-shot answer, p0, and K_max verified candidates with their confidences (to cache_dir too,
if given). Each Optuna TPE trial only re-selects demos from the cache with the trial's ABA/DF rule
and re-answers. Answers are memoized per (query, selected subset), so later trials are almost free.
Only a few demo subsets per query are reachable under any ABA/DF setting (≤ 15 for K_max=4); they are
answered in one batch up front, so trials are lookups and hundreds of them take seconds. The search
covers epsilon, A0, A1, B, t_low, t_high (and k_max with k_max_choices, e.g. [4, 6] as in the
paper). The paper defaults and plain K_max-shot ICL are evaluated first. If no configuration beats
zero-shot on the validation set, the tuned config never imagines (zero-shot answers, no extra cost).
The paper uses 50 validation examples per benchmark.
Runtime budget, which matters because synthesis dominates runtime:
max_runtime_ratio=25caps the estimated batched runtime relative to zero-shot. The estimate comes from three timings measured on your machine during tuning: the batched zero-shot time per query, the batched synthesis time per candidate (saved ascost.jsonincache_dirfor later runs), and the ICL answer time. In our runs the measured test-time ratio was within about ±25% of the estimate, so leave some margin.max_candidatescaps the mean number of candidates synthesized per query.1.0withK_max=4is the paper's operating point (ABA synthesizes ~23% ofK_max).- With neither, the search maximizes accuracy alone and usually spends the whole budget.
How much one candidate costs relative to a zero-shot answer depends heavily on the model and task. With BAGEL, a candidate needs ~1.3-1.9 generated images (critic retries), each taking ~2 s on an H100. A short VQA answer takes ~0.1 s.
result = model.tune(val, metric="vqa_accuracy", n_trials=300, max_runtime_ratio=25, cache_dir="cache/vizwiz")
print(result) # best/zero-shot score, gain, candidates/demos per query, est. runtime, params
model.config.save("vizwiz.json")
model.config = SIMITConfig.load("vizwiz.json")
Custom skills
A skill turns a text description into an image through a model-written spec. Subclass
simit.Skill: the router learns the new category from route_hint and examples. If parse or
render raises SpecError(msg), msg is shown to the model so it can repair its spec.
import simit
from simit import SpecError
class SheetMusic(simit.Skill):
name = "sheet_music"
route_hint = "musical notation: staves, notes, chords"
examples = [ # (description, spec) few-shot pairs for the spec prompt
("A C major scale in quarter notes.", "X:1\nK:C\nL:1/4\nCDEF GABc|"),
("A G major chord held for a whole note.", "X:1\nK:G\nL:1\n[GBd]|"),
]
def parse(self, text):
abc = simit.extract_fenced_block(text) or text
if "K:" not in abc:
raise SpecError("the ABC spec needs a key line such as 'K:C'")
return abc
def render(self, abc) -> "PIL.Image.Image":
return my_abc_renderer(abc) # any deterministic renderer
model.add_skill(SheetMusic())
Optional attributes and hooks:
route_examples: example descriptions for the router (default: the descriptions inexamples).prompt: a full spec prompt with{request}to replace the one built fromexamples.max_new_tokens,first_temperature,retry_temperature.verify_rubric = "strict" | "structured".render_in_subprocess: renders run in a crash- and timeout-isolated worker pool. Skills defined in__main__or holding unpicklable state run in a thread instead.shortcut(request): render without a model call.clean(text).
The 17 built-in skills: natural, diagram, graph, molecule, circuit, vector, venn, table, scene, puzzle, geometry, figure (sandboxed matplotlib), mermaid, svg, vegalite, html, html_composite (multi-panel layouts whose panels are realized recursively).
Loading once, running per request (e.g. Hugging Face ZeroGPU)
Model weights can be loaded once and wrapped in a fresh engine per request, which is what a ZeroGPU Space needs (it forks the process for every GPU call, so no engine threads may exist in the main process, and the main process must not initialize CUDA):
from simit import SIMIT, SIMITConfig
from simit.backends.bagel import BagelBackend, BagelModel
weights = BagelModel.load("ByteDance-Seed/BAGEL-7B-MoT", device="cpu") # at startup
@spaces.GPU(duration=90)
def run(image, question):
sim = SIMIT(BagelBackend(weights, device="cuda", use_cuda_graphs=False), config=SIMITConfig(k_max=2))
try:
demos = sim.imagine(image, question, time_limit=60)
return sim.greedy(image, question), sim.answer(image, question, demos)
finally:
sim.close()
LanceModel.load / LanceBackend work the same way. For a transformers model, pass the loaded
model and processor to HFBackend. The SIMIT demo Space (three models, per-budget presets, a
streaming UI) is built this way.
New model families
simit.backends.register_backend(name, detect, factory) adds a backend. It subclasses
simit.backends.Backend and implements submit_generate, submit_score and, for native image
generation, submit_image. pipeline_defaults can change synthesis defaults for that model
(see LanceBackend).
Configuration
SIMITConfig (model.config, or from_pretrained(..., config=...)). None means the backend's default.
| field | default | meaning |
|---|---|---|
k_max |
4 | max demos per query |
use_aba, epsilon, A0, A1, B |
True, 0.14, 0.35, 0.60, 0.2 | adaptive budget allocation (App. D) |
use_df, t_low, t_high |
True, 0.2, 0.9 | difficulty-filter confidence band |
synthesis |
"batch" (Lance: "decomposed") |
one call for all triplets vs. three short calls per triplet |
diversity_prompt |
True | the paper's diversity-encouraging triplet prompt |
use_skills |
True (Lance: False) | structured skill library vs. natural images only |
verify, verify_threshold, verify_think |
True (Lance: False), 50, BAGEL: True | critic |
attempts_per_slot, repair_retries, verify_rounds |
4, 2, 5 | breadth-first slot-filling budget |
image_size |
400 | natively generated image side |
speculative |
on for imagine, off for imagine_batch/tune |
start all remaining candidate attempts at once (lower latency for a lone query; wasted work when the GPU is busy) |
answer_max_new_tokens |
128 | answer length (also per call: max_new_tokens=) |
seed |
None | image-generation seed |
Environment variables:
SIMIT_CHROMIUM: path to a Chromium/Chrome binary.PLAYWRIGHT_BROWSERS_PATH: where Playwright's browsers live.SIMIT_RENDER_WORKERS: number of render processes. The default is derived from the CPUs available to the job (LSF/SLURM aware).
Performance
Everything that can overlap does:
- Built-in engine for BAGEL/Lance, running both the understanding and generation experts:
- continuous batching of decode, prefill and diffusion steps from all concurrent requests
- automatic prefix caching of shared prompt prefixes (the long router and skill few-shot prompts are prefilled once)
- CUDA graphs for decode
- experts run on contiguous token slices, with fused Triton kernels for RMSNorm, RoPE and SwiGLU (bit-exact with the eager ops up to reduction order)
- separate high-priority text and image streams, so diffusion doesn't stall decoding
- vLLM for standard VLMs (transformers fallback batches concurrent requests).
- Pipeline:
- triplets are streamed into realization as they are decoded, and synthesis stops once enough are kept
- breadth-first slot filling: every triplet gets one cheap attempt before failures get repairs
- the online ABA/DF rule stops as soon as
K*demos pass. A loneimaginecall runs the remaining candidate attempts in parallel (speculative), since an idle GPU makes that nearly free - renders run in a process pool (warm matplotlib zygote, one shared headless browser)
imagine_batch/tunerun many queries concurrently so their calls share GPU batches.
Measured on one H100 with tests/bench.py: LMMs-Eval-Lite slices, tuned on 60 validation queries
(300 trials), tested on 60 different queries. Runtime is the batched wall time relative to batched
zero-shot answering. With 60 test queries, differences of a few points are within noise.
Tuned with max_runtime_ratio=25:
| model | task | test zero-shot → SIMIT-ICL | runtime (estimated) |
|---|---|---|---|
| BAGEL-7B-MoT | VizWiz | 0.494 → 0.556 (+12.4%) | 25.0× (22.0×) |
| BAGEL-7B-MoT | OK-VQA | 0.650 → 0.650 (nothing beat zero-shot on validation) | 1.0× |
| Lance | VizWiz | 0.106 → 0.089 | 27.0× (21.3×) |
| Lance | OK-VQA | 0.483 → 0.489 (+1.1%) | 17.4× (23.3×) |
Tuned with max_candidates=1.0:
| model | task | test zero-shot → SIMIT-ICL | runtime |
|---|---|---|---|
| BAGEL-7B-MoT | VizWiz | 0.494 → 0.567 (+14.6%) | 29.3× |
| BAGEL-7B-MoT | OK-VQA | 0.650 → 0.661 (+1.7%) | 64.6× |
| BAGEL-7B-MoT | ChartQA | 0.817 → 0.817 (zero-shot chosen) | 1.0× |
| BAGEL-7B-MoT | AI2D | 0.950 → 0.933 | 12.1× |
| Qwen3.8-27B + FLUX.2-klein-4B (vLLM) | VizWiz | 0.522 → 0.522 | 12.0× |
| Qwen3.8-27B + FLUX.2-klein-4B (vLLM) | OK-VQA / ChartQA | unchanged (zero-shot chosen) | 0.9× |
| Qwen3.8-27B + FLUX.2-klein-4B (vLLM) | AI2D | 0.900 → 0.900 | 49.7× |
What the numbers show:
- BAGEL's VizWiz gain reproduces the paper's direction.
- Lance matches the research implementation's VizWiz behavior. Its zero-shot answers are mostly "No", which gives a very low baseline.
- Qwen3.8-27B gains little from imagined demos on these tasks.
- The OK-VQA and AI2D rows at 50-65× show why a runtime cap beats a candidate cap. One BAGEL candidate needs ~1.3-1.9 generated images at ~2 s each (50 steps, CFG), while a short answer takes ~0.1 s.
Single-query latency, i.e. what one demo request feels like:
- BAGEL: 13.9 s on average on VizWiz (16 s for queries that get demos), with speculative attempts. Greedy takes 0.29 s.
- Qwen3.8-27B + FLUX: ~20 s for
imagine(k=2). - Lance: ~1-2 s.
The BAGEL/Lance engine matches the research implementation: 17/18 (BAGEL) and 15/18 (Lance) identical greedy answers on reference samples, with the rest near-ties.
Differences from the research code
- The tuner caches
K_maxcandidates and memoizes answers per demo subset (as in App. G), with an optional synthesis budget (max_candidates). - Fixed a cleanup bug that stripped valid CSS after CJK characters in HTML specs.
- Duplicate questions within a query's demos are dropped, and a
[...]copied from the triplet template around a whole field (Answer: [3]) is removed, as are chat preambles ("Sure, here's the first one: ...") in decomposed questions. - Triplet synthesis is streamed and stops early.
- When the critic's reasoning hits the token limit before its
SCORE:line, a short follow-up asks for the score instead of discarding the candidate. - Prompts (triplet, router, skills, critic, revision, ICL format) are the research prompts verbatim; the triplet prompt is question-free, as in the code.
License
Apache-2.0. simit/backends/mot/wan_vae.py is adapted from vllm-omni (Apache-2.0). Mermaid
(MIT) is bundled under simit/skills/assets/.
Metadata
Release files for simit 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| simit-0.1.1.tar.gz | 1.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| simit-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.3 MB
Release files / simit-0.1.1.tar.gz
| Download URL | simit-0.1.1.tar.gz |
|---|---|
| Size | 1.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
947a629b71a1e5ecf2510501051833681a560719313e95f57ff4cd9456c6cf2d
|
|
BLAKE2b-256 checksum How to use checksums |
a14187648823ef6ba2962d05c2c0b07656a4bf6891c9d0d1ffeee844717bcc4b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / simit-0.1.1-py3-none-any.whl
| Download URL | simit-0.1.1-py3-none-any.whl |
|---|---|
| Size | 1.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2dc6fff444d4f7c241c80623a109b587cc94480f9e5f94f0f6b67363498c5bad
|
|
BLAKE2b-256 checksum How to use checksums |
8cad6ee98d7403c2edb15f134a38f8a34473ef9b5e3b8838ba05962ee9a9867d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log