Skip to main content

SIMIT-ICL

Vision-language models that imagine their own in-context demonstrations at test time.

Given an unlabeled query (image, question), SIMIT-ICL has the model synthesize a few similar (image, question, answer) examples, verifies them, keeps only the useful ones, and answers the query with them in context. No labels, no extra training, and one model does all the work (or a standard VLM plus a text-to-image tool).

from simit import SIMIT

model = SIMIT.from_pretrained("ByteDance-Seed/BAGEL-7B-MoT")

demos  = model.imagine(image, question)          # imagined (image, question, answer) demos
answer = model.answer(image, question, demos)    # SIMIT-ICL answer
greedy = model.greedy(image, question)           # standard zero-shot answer, for comparison

How one query is processed (paper, Sec. 3 / App. G):

  1. Zero-shot pass: greedy answer + confidence p0 (geometric-mean token probability).
  2. Adaptive budget (ABA): K*(p0) demos, 0 for queries the model is already sure about.
  3. Triplet synthesis: the model proposes (description, question, answer) triplets similar to the query.
  4. Realization: a router picks one of 17 skills. Natural images come from the model's own image head (BAGEL, Lance) or an image-generator tool (standard VLMs). Charts, diagrams, molecules, circuits, tables, flowcharts, etc. are written as code/specs by the model and rendered deterministically, with parse/render errors fed back for repair.
  5. Critic: the model scores each image against its description (0-100). Failures are revised and regenerated.
  6. Difficulty filter (DF): keep demos whose teacher-forced answer confidence lies in [t_low, t_high].
  7. ICL answer with the kept demos in context.

Installation

pip install simit              # BAGEL, Lance, any transformers VLM
pip install "simit[vllm]"      # + the vLLM engine for standard VLMs (much faster than plain transformers)

# or from source
git clone https://github.com/monurcan/simit && cd simit && pip install -e .

Every skill dependency is a pip wheel: there is no Node.js, Mermaid CLI, cairo or manual browser setup. HTML and Mermaid skills render in headless Chromium through Playwright, with the following fallbacks:

  1. Playwright's own browser, if one is installed.
  2. A system chromium/chrome.
  3. Otherwise SIMIT downloads Playwright's headless Chromium once on first use (~100 MB, file-locked so parallel processes don't race). If Playwright's downloader fails, as it does behind some HPC networks, it fetches the same build with plain Python (proxy environment variables are honored).

Mermaid itself ships inside the package.

Requirements: Python ≥ 3.10, PyTorch ≥ 2.12 (BAGEL/Lance use torch.nn.attention.varlen, so no flash-attn build is needed), one CUDA GPU (H100/A100 80GB for the 27B example; BAGEL-7B and Lance fit on smaller cards).

Supported models

model how images are made engine
ByteDance-Seed/BAGEL-7B-MoT native image head + 17 skills, critic in thinking mode (as in the paper) built-in MoT engine
bytedance-research/Lance native image head; defaults to decomposed synthesis, natural images, no critic (its instruction following is weaker) built-in MoT engine
any transformers VLM, e.g. Qwen/Qwen3.8-27B image_generator= tool (e.g. black-forest-labs/FLUX.2-klein-4B) + 17 skills; omit the tool for skills only vLLM if installed, else transformers
# a standard VLM + a text-to-image tool
model = SIMIT.from_pretrained("Qwen/Qwen3.8-27B", image_generator="black-forest-labs/FLUX.2-klein-4B")

# force an engine
model = SIMIT.from_pretrained("Qwen/Qwen3.8-27B", engine="transformers")   # or "vllm"

# wrap a model you already loaded with the standard HF interface
from transformers import AutoProcessor, AutoModelForMultimodalLM
processor = AutoProcessor.from_pretrained("Qwen/Qwen3.8-27B")
hf_model = AutoModelForMultimodalLM.from_pretrained("Qwen/Qwen3.8-27B", device_map="auto")
model = SIMIT.from_pretrained(hf_model, processor=processor, image_generator="black-forest-labs/FLUX.2-klein-4B")

image_generator can be a diffusers model id, a loaded diffusers pipeline, or any callable (prompt, width, height) -> PIL.Image. On one GPU it is loaded first and the VLM engine sizes its memory around it. With several GPUs it goes on the last one.

Usage

from simit import SIMIT, save_demos

model = SIMIT.from_pretrained("ByteDance-Seed/BAGEL-7B-MoT")

# 1) fit ABA + DF on a small labeled validation set (synthesis runs once and is cached)
model.tune(val_set, metric="vqa_accuracy", max_runtime_ratio=25)

# 2) test time
for image, question in test_set:
    demos = model.imagine(image, question)
    improved = model.answer(image, question, demos)

Main methods:

method what it does
imagine(image, question, k=None, return_details=False, time_limit=None, on_demo=None, on_progress=None) adaptive number of demos (ABA + DF); k=n asks for n verified demos (fewer if the attempt budget runs out first, e.g. when the critic keeps rejecting images); return_details=True returns an Imagination with the zero-shot answer, p0, the budget, every candidate and stats. time_limit (seconds) stops and returns the demos accepted so far; on_demo(demo) / on_progress(stage) stream progress, e.g. to a UI
zero_shot(image, question) the greedy answer with its confidence p0; model.config.budget(zs.confidence) is the number of demos ABA will ask for. Reused by a following imagine
answer(image, question, demos) answer with the demos in context (no demos: zero-shot)
greedy(image, question) the standard zero-shot greedy answer (cached and shared with imagine)
model(image, question) imagine then answer
imagine_batch(queries) / answer_batch(queries, demos) many queries at once; their model calls are batched together on the GPU (much faster than a loop)
tune(val_set, metric, n_trials=100, max_runtime_ratio=None, max_candidates=None, cache_dir=None, k_max_choices=None) hyperparameter search (below)
add_skill(skill) / remove_skill(name) change the skill library

val_set items are (image, question, answer_or_answers) tuples or dicts with image/question/answer(s) (and optionally their own metric, for validation sets that mix benchmarks). metric is one of exact_match, vqa_accuracy, contains, anls, relaxed_accuracy, multiple_choice, or any fn(prediction, references) -> float. Images can be PIL images, file paths or raw bytes.

Saving demos for fine-tuning (SIMIT-FT)

from simit import save_demos, load_demos
save_demos(demos, "imagined/query_0")      # PNGs + demos.jsonl (with LLaVA-style "conversations")

The JSONL has image, question, answer, description, skill, confidence, verify_score and a conversations field, ready for LLaVA-format fine-tuning scripts. SIMIT-FT itself (fine-tuning on the imagined test-set data) is out of scope for this package.

Hyperparameter tuning

tune follows App. G. For each validation query it runs synthesis once and caches the zero-shot answer, p0, and K_max verified candidates with their confidences (to cache_dir too, if given). Each Optuna TPE trial only re-selects demos from the cache with the trial's ABA/DF rule and re-answers. Answers are memoized per (query, selected subset), so later trials are almost free. Only a few demo subsets per query are reachable under any ABA/DF setting (≤ 15 for K_max=4); they are answered in one batch up front, so trials are lookups and hundreds of them take seconds. The search covers epsilon, A0, A1, B, t_low, t_high (and k_max with k_max_choices, e.g. [4, 6] as in the paper). The paper defaults and plain K_max-shot ICL are evaluated first. If no configuration beats zero-shot on the validation set, the tuned config never imagines (zero-shot answers, no extra cost). The paper uses 50 validation examples per benchmark.

Runtime budget, which matters because synthesis dominates runtime:

  • max_runtime_ratio=25 caps the estimated batched runtime relative to zero-shot. The estimate comes from three timings measured on your machine during tuning: the batched zero-shot time per query, the batched synthesis time per candidate (saved as cost.json in cache_dir for later runs), and the ICL answer time. In our runs the measured test-time ratio was within about ±25% of the estimate, so leave some margin.
  • max_candidates caps the mean number of candidates synthesized per query. 1.0 with K_max=4 is the paper's operating point (ABA synthesizes ~23% of K_max).
  • With neither, the search maximizes accuracy alone and usually spends the whole budget.

How much one candidate costs relative to a zero-shot answer depends heavily on the model and task. With BAGEL, a candidate needs ~1.3-1.9 generated images (critic retries), each taking ~2 s on an H100. A short VQA answer takes ~0.1 s.

result = model.tune(val, metric="vqa_accuracy", n_trials=300, max_runtime_ratio=25, cache_dir="cache/vizwiz")
print(result)                 # best/zero-shot score, gain, candidates/demos per query, est. runtime, params
model.config.save("vizwiz.json")
model.config = SIMITConfig.load("vizwiz.json")

Custom skills

A skill turns a text description into an image through a model-written spec. Subclass simit.Skill: the router learns the new category from route_hint and examples. If parse or render raises SpecError(msg), msg is shown to the model so it can repair its spec.

import simit
from simit import SpecError

class SheetMusic(simit.Skill):
    name = "sheet_music"
    route_hint = "musical notation: staves, notes, chords"
    examples = [                                  # (description, spec) few-shot pairs for the spec prompt
        ("A C major scale in quarter notes.", "X:1\nK:C\nL:1/4\nCDEF GABc|"),
        ("A G major chord held for a whole note.", "X:1\nK:G\nL:1\n[GBd]|"),
    ]

    def parse(self, text):
        abc = simit.extract_fenced_block(text) or text
        if "K:" not in abc:
            raise SpecError("the ABC spec needs a key line such as 'K:C'")
        return abc

    def render(self, abc) -> "PIL.Image.Image":
        return my_abc_renderer(abc)               # any deterministic renderer

model.add_skill(SheetMusic())

Optional attributes and hooks:

  • route_examples: example descriptions for the router (default: the descriptions in examples).
  • prompt: a full spec prompt with {request} to replace the one built from examples.
  • max_new_tokens, first_temperature, retry_temperature.
  • verify_rubric = "strict" | "structured".
  • render_in_subprocess: renders run in a crash- and timeout-isolated worker pool. Skills defined in __main__ or holding unpicklable state run in a thread instead.
  • shortcut(request): render without a model call.
  • clean(text).

The 17 built-in skills: natural, diagram, graph, molecule, circuit, vector, venn, table, scene, puzzle, geometry, figure (sandboxed matplotlib), mermaid, svg, vegalite, html, html_composite (multi-panel layouts whose panels are realized recursively).

Loading once, running per request (e.g. Hugging Face ZeroGPU)

Model weights can be loaded once and wrapped in a fresh engine per request, which is what a ZeroGPU Space needs (it forks the process for every GPU call, so no engine threads may exist in the main process, and the main process must not initialize CUDA):

from simit import SIMIT, SIMITConfig
from simit.backends.bagel import BagelBackend, BagelModel

weights = BagelModel.load("ByteDance-Seed/BAGEL-7B-MoT", device="cpu")   # at startup

@spaces.GPU(duration=90)
def run(image, question):
    sim = SIMIT(BagelBackend(weights, device="cuda", use_cuda_graphs=False), config=SIMITConfig(k_max=2))
    try:
        demos = sim.imagine(image, question, time_limit=60)
        return sim.greedy(image, question), sim.answer(image, question, demos)
    finally:
        sim.close()

LanceModel.load / LanceBackend work the same way. For a transformers model, pass the loaded model and processor to HFBackend. The SIMIT demo Space (three models, per-budget presets, a streaming UI) is built this way.

New model families

simit.backends.register_backend(name, detect, factory) adds a backend. It subclasses simit.backends.Backend and implements submit_generate, submit_score and, for native image generation, submit_image. pipeline_defaults can change synthesis defaults for that model (see LanceBackend).

Configuration

SIMITConfig (model.config, or from_pretrained(..., config=...)). None means the backend's default.

field default meaning
k_max 4 max demos per query
use_aba, epsilon, A0, A1, B True, 0.14, 0.35, 0.60, 0.2 adaptive budget allocation (App. D)
use_df, t_low, t_high True, 0.2, 0.9 difficulty-filter confidence band
synthesis "batch" (Lance: "decomposed") one call for all triplets vs. three short calls per triplet
diversity_prompt True the paper's diversity-encouraging triplet prompt
use_skills True (Lance: False) structured skill library vs. natural images only
verify, verify_threshold, verify_think True (Lance: False), 50, BAGEL: True critic
attempts_per_slot, repair_retries, verify_rounds 4, 2, 5 breadth-first slot-filling budget
image_size 400 natively generated image side
speculative on for imagine, off for imagine_batch/tune start all remaining candidate attempts at once (lower latency for a lone query; wasted work when the GPU is busy)
answer_max_new_tokens 128 answer length (also per call: max_new_tokens=)
seed None image-generation seed

Environment variables:

  • SIMIT_CHROMIUM: path to a Chromium/Chrome binary.
  • PLAYWRIGHT_BROWSERS_PATH: where Playwright's browsers live.
  • SIMIT_RENDER_WORKERS: number of render processes. The default is derived from the CPUs available to the job (LSF/SLURM aware).

Performance

Everything that can overlap does:

  • Built-in engine for BAGEL/Lance, running both the understanding and generation experts:
    • continuous batching of decode, prefill and diffusion steps from all concurrent requests
    • automatic prefix caching of shared prompt prefixes (the long router and skill few-shot prompts are prefilled once)
    • CUDA graphs for decode
    • experts run on contiguous token slices, with fused Triton kernels for RMSNorm, RoPE and SwiGLU (bit-exact with the eager ops up to reduction order)
    • separate high-priority text and image streams, so diffusion doesn't stall decoding
  • vLLM for standard VLMs (transformers fallback batches concurrent requests).
  • Pipeline:
    • triplets are streamed into realization as they are decoded, and synthesis stops once enough are kept
    • breadth-first slot filling: every triplet gets one cheap attempt before failures get repairs
    • the online ABA/DF rule stops as soon as K* demos pass. A lone imagine call runs the remaining candidate attempts in parallel (speculative), since an idle GPU makes that nearly free
    • renders run in a process pool (warm matplotlib zygote, one shared headless browser)
  • imagine_batch/tune run many queries concurrently so their calls share GPU batches.

Measured on one H100 with tests/bench.py: LMMs-Eval-Lite slices, tuned on 60 validation queries (300 trials), tested on 60 different queries. Runtime is the batched wall time relative to batched zero-shot answering. With 60 test queries, differences of a few points are within noise.

Tuned with max_runtime_ratio=25:

model task test zero-shot → SIMIT-ICL runtime (estimated)
BAGEL-7B-MoT VizWiz 0.494 → 0.556 (+12.4%) 25.0× (22.0×)
BAGEL-7B-MoT OK-VQA 0.650 → 0.650 (nothing beat zero-shot on validation) 1.0×
Lance VizWiz 0.106 → 0.089 27.0× (21.3×)
Lance OK-VQA 0.483 → 0.489 (+1.1%) 17.4× (23.3×)

Tuned with max_candidates=1.0:

model task test zero-shot → SIMIT-ICL runtime
BAGEL-7B-MoT VizWiz 0.494 → 0.567 (+14.6%) 29.3×
BAGEL-7B-MoT OK-VQA 0.650 → 0.661 (+1.7%) 64.6×
BAGEL-7B-MoT ChartQA 0.817 → 0.817 (zero-shot chosen) 1.0×
BAGEL-7B-MoT AI2D 0.950 → 0.933 12.1×
Qwen3.8-27B + FLUX.2-klein-4B (vLLM) VizWiz 0.522 → 0.522 12.0×
Qwen3.8-27B + FLUX.2-klein-4B (vLLM) OK-VQA / ChartQA unchanged (zero-shot chosen) 0.9×
Qwen3.8-27B + FLUX.2-klein-4B (vLLM) AI2D 0.900 → 0.900 49.7×

What the numbers show:

  • BAGEL's VizWiz gain reproduces the paper's direction.
  • Lance matches the research implementation's VizWiz behavior. Its zero-shot answers are mostly "No", which gives a very low baseline.
  • Qwen3.8-27B gains little from imagined demos on these tasks.
  • The OK-VQA and AI2D rows at 50-65× show why a runtime cap beats a candidate cap. One BAGEL candidate needs ~1.3-1.9 generated images at ~2 s each (50 steps, CFG), while a short answer takes ~0.1 s.

Single-query latency, i.e. what one demo request feels like:

  • BAGEL: 13.9 s on average on VizWiz (16 s for queries that get demos), with speculative attempts. Greedy takes 0.29 s.
  • Qwen3.8-27B + FLUX: ~20 s for imagine(k=2).
  • Lance: ~1-2 s.

The BAGEL/Lance engine matches the research implementation: 17/18 (BAGEL) and 15/18 (Lance) identical greedy answers on reference samples, with the rest near-ties.

Differences from the research code

  • The tuner caches K_max candidates and memoizes answers per demo subset (as in App. G), with an optional synthesis budget (max_candidates).
  • Fixed a cleanup bug that stripped valid CSS after CJK characters in HTML specs.
  • Duplicate questions within a query's demos are dropped, and a [...] copied from the triplet template around a whole field (Answer: [3]) is removed, as are chat preambles ("Sure, here's the first one: ...") in decomposed questions.
  • Triplet synthesis is streamed and stops early.
  • When the critic's reasoning hits the token limit before its SCORE: line, a short follow-up asks for the score instead of discarding the candidate.
  • Prompts (triplet, router, skills, critic, revision, ICL format) are the research prompts verbatim; the triplet prompt is question-free, as in the code.

License

Apache-2.0. simit/backends/mot/wan_vae.py is adapted from vllm-omni (Apache-2.0). Mermaid (MIT) is bundled under simit/skills/assets/.

Metadata

Release files for simit 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for simit 0.1.1
File Size Uploaded
simit-0.1.1.tar.gz 1.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for simit 0.1.1
File Interpreter ABI Platform
simit-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 2.3 MB

Release files / simit-0.1.1.tar.gz

Download URL simit-0.1.1.tar.gz
Size 1.1 MB
Tags Source
SHA-256 checksum
How to use checksums
947a629b71a1e5ecf2510501051833681a560719313e95f57ff4cd9456c6cf2d
BLAKE2b-256 checksum
How to use checksums
a14187648823ef6ba2962d05c2c0b07656a4bf6891c9d0d1ffeee844717bcc4b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / simit-0.1.1-py3-none-any.whl

Download URL simit-0.1.1-py3-none-any.whl
Size 1.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
2dc6fff444d4f7c241c80623a109b587cc94480f9e5f94f0f6b67363498c5bad
BLAKE2b-256 checksum
How to use checksums
8cad6ee98d7403c2edb15f134a38f8a34473ef9b5e3b8838ba05962ee9a9867d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page