Skip to main content

Vega: a particle at rest in the deeper of two wells

vegaml

Typed decisions from a frozen language model and a small physics engine. You give it a state and a question whose answer type is fixed in advance; it returns a value your code can use directly: a choice, a 0 to 1 score, or the probability that a statement is true. Each carries a calibrated probability, a conformal answer set and an explicit abstain flag.

800M or 4B parameters · 73,728-token context · images · runs on your own hardware. No text is generated anywhere in the path.

Where this model beats a hosted reference

pip install vegaml

Open in Colab

Inference on a single T4, end to end, in the browser.

Releases are cut by tagging vX.Y.Z: CI builds, checks the version matches the tag, and publishes through PyPI Trusted Publishing, so no API token exists in this repository.

Quick start

import vegaml

v = vegaml.load()            # the 0.8B, mode="engine": both are the defaults
# v = vegaml.load("4b")      # the 4B, same repository, same interface

out = v.decide(
    {"from": "billing@acme.com", "subject": "Invoice overdue", "body": "Third notice. Pay now."},
    {"team":  {"type": "choice", "instructions": "Which team should handle this?",
               "criteria": {"billing": "payments and invoices",
                            "technical": "product faults",
                            "sales": "new business"}},
     "churn": {"type": "boolean", "instructions": "Is this customer at risk of leaving?",
               "criteria": {"true": "shows intent to cancel", "false": "no such signal"}}})

print(out["answers"]["team"]["choice"])        # 'billing'
print(out["answers"]["team"]["probs"])         # calibrated, sums to 1
print(out["answers"]["churn"]["abstain"])      # False -> the engine stands behind it

Both questions are answered from one read of the state.

Two readouts, one set of features

mode picks how the pooled features are turned into an answer.

mode needs examples carries use it when
"engine" (default) no, zero-shot calibration, conformal sets, abstain the normal case, and every published figure below
"ttt" yes, 6 minimum nothing, and it answers even when it should not you have labels for exactly this question
"both" yes both, side by side deciding which to trust
examples = [({"body": "cancel my account"}, {"churn": "true"}),
            ({"body": "how do I export?"},  {"churn": "false"})]            # ... 20 or so

report = v.fit(examples, questions)
print(report["churn"]["cv_accuracy"], report["churn"]["at_or_below_chance"])

out = v.decide(state, questions, mode="ttt")

fit returns a cross-validated accuracy per question. Read it. A head at or below chance has learned nothing and its probabilities are noise; the engine is the honest answer there.

Images

from PIL import Image

v.decide_image(Image.open("invoice.png"),
               {"kind": {"type": "choice", "instructions": "What kind of document is this?",
                         "criteria": {"invoice": "a bill", "receipt": "proof of payment",
                                      "form": "a document to be filled in", "other": "none of these"}}})

The backbone is multimodal and its vision encoder is frozen with the rest, so a picture takes the place of the state text in the same prompt and the same spans are pooled. The benchmark figures below are text; the image path is functional but not covered by them.

73,728-token context

max_len is 73,728 tokens, and the architecture is built for it rather than merely permitting it:

  • Read-once prefix caching. A state of at least 4,096 tokens is encoded once; every question about that state continues from a copy of that prefix. Twelve questions about one 73k-token contract cost roughly one read, not twelve.
  • A memory reader. Long states are cut into 16 pieces, mean-pooled at the deepest feature layer and fed to the engine alongside the spans, so evidence late in a long document still reaches the decision.
  • No silent truncation. An input that will not fit is refused, never quietly shortened.

Answer types

Three, named in every question's type field:

type returns
choice one option from a set you define
score an integer level against ordered criteria, plus the expected value
boolean the probability that a statement holds

Changed in 0.5.0. boolean was renamed from a coined word that told a reader nothing. A question using the old name is rejected with the valid types listed, rather than silently reinterpreted. Checkpoints published before the change key their conformal thresholds by the old name; the loader remaps them, so an older checkpoint still loads with its prediction sets intact.

Two sizes

Both live in the same checkpoint repository and take the same code path. load() with no argument gives you the 0.8B, which is the default and the model every Jev comparison on this page measures.

0.8B 4B
backbone, frozen Qwen3.5-0.8B Qwen3.5-4B
engine, trained 14.3M params, 57 MB 30.9M params, 124 MB
task adapter 0.84M params 1.43M params
with the task adapter 0.763 0.803
soft accuracy 0.680 0.725
calibration error 0.026 0.019
per decision 22 ms 28 ms

Measured on the LocalLLaMA/typed-decisions test split, 2,050 decisions, each checkpoint's own run of the same harness. Per workflow with the adapter, the 4B leads on five of six: apibank 96.5 against 89.5, when2call 94.6 against 90.4, appliances 89.1 against 85.1, banking77 74.9 against 71.4, esci 64.4 against 56.4. The 0.8B wins clinc150, 88.0 against 86.7.

So the 4B is the better model and the 0.8B is the default anyway, because 22 ms against 28 ms and 57 MB against 124 MB decides more deployments than four points of accuracy does.

vegaml.load()                  # 0.8B
vegaml.load("4b")              # 4B
vegaml.load("owner/some-repo") # anything else on the Hub
vegaml.load("/path/on/disk")   # or a local directory

A private or gated repository needs a Hub token, and without one the Hub answers 401, which it reports identically to "no such repository". fetch turns that into one sentence naming which of the two it is and where it looked for a token. load() picks one up from HF_TOKEN in the environment, so nothing has to be threaded through the call; pass token=... to override it, or token=False to ignore the environment and fetch anonymously. It stays optional: the published checkpoints are public and load with no token set.

vegaml.load("owner/private-repo")               # uses $HF_TOKEN if it is set
vegaml.load("owner/private-repo", token=tok)    # or pass it explicitly

Fine-tuning

notebooks/vega_finetune_2xT4_kaggle.ipynb trains a gated adapter end to end on Kaggle's 2×T4, and notebooks/vega_finetune.py is the script it writes to disk and runs, so the same three commands work on any machine with a GPU.

The language model stays frozen. What is trained is a rank-32 adapter on seven engine projections plus its gate, about 0.84M parameters, on dair-ai/emotion as a typed choice question, with gate negatives drawn from ag_news so the adapter learns when not to fire. An adapter that always fires is not routable. Thermal calibration and the conformal thresholds are refitted afterwards on a validation split that training never sees, and travel with the adapter.

The three stages

pip install vegaml datasets

# 1. frozen features, sharded across both GPUs, cached to disk. The long part.
torchrun --nproc_per_node=2 vega_finetune.py extract --batch 32

# 2. the adapter, on those cached features. Minutes.
python vega_finetune.py train --epochs 3 --batch 64 --rank 32 --lr 3e-4

# 3. base engine against the adapter on the held-out split
python vega_finetune.py eval --batch 64

Only stage 1 uses both GPUs, because only stage 1 benefits. The frozen read happens exactly once per row and is the expensive part; training then runs on cached features, where an epoch is seconds on one GPU. Launching stages 2 and 3 under torchrun would buy nothing, so they are not.

Settings

variable default what it does
VEGA_SIZE 0.8b which size to tune, 0.8b or 4b
HF_TOKEN unset needed for a private or gated checkpoint, and to push
VEGA_MODEL_REPO the size name a different checkpoint repository or a local path
VEGA_BACKBONE per size a different frozen backbone
VEGA_WORK /kaggle/working where features, the adapter and results are written

The two backbones pool to different widths, so each size keeps its own feature cache (features-0.8b, features-4b) and its own output directory. Switching VEGA_SIZE never mixes the two or silently reads the wrong cache.

In the notebook these live in one settings cell near the top, which also reads HF_TOKEN from Kaggle Secrets when you have added one and carries on when you have not.

Pushing what you trained

export HF_TOKEN=...
python vega_finetune.py eval --batch 64 --push-to your-name/vega-emotion-adapter

Only the adapter and its metadata go up, a few hundred kilobytes: no backbone and no engine. The repository is created private unless you add --public. With no token the push is skipped with a logged line rather than failing the run. VEGA_PUSH_TO sets the destination if you would rather not pass the flag.

Using the result

adapter-<size>/ holds emotion.safetensors and emotion.json in the layout the loader already reads. Drop both into a checkpoint's adapters/ directory and they load automatically. The gate decides per question whether they apply, and the adapter's own calibration travels with it.

import vegaml

v = vegaml.load("your-copy-of-the-checkpoint")   # adapters/ picked up automatically
v.decide("I can't believe they remembered my birthday",
         {"emotion": {"type": "choice",
                      "instructions": "Which emotion does this message express?",
                      "criteria": {"sadness": "...", "joy": "...", "love": "...",
                                   "anger": "...", "fear": "...", "surprise": "..."}}})

The adapter records the SHA-256 of the base engine it was trained against, and the loader refuses it against a different one rather than silently producing nonsense.

What has actually been run

All three stages were run end to end before this was committed, but on a laptop and a 96-row slice, not on two T4s: extract → train → eval completed, two-rank sharding was checked to recombine each split exactly once, and the adapter moved test accuracy 0.313 → 0.542 on 48 held-out items. Those numbers are a smoke test, not a result. The full 16k run on Kaggle is the real one.

Benchmarks

The figure at the top of this page is the table below. Item-paired against a live Jev 1.13.0 over identical items, no tuning on the evaluation data. These are the areas where this model leads; the reference leads on most others, particularly reranking and multi-step reasoning.

task this model Jev 1.13.0
Phishing screening, 800 emails 75.4 acc · 252/400 caught · recall 0.63 61.9 · 99/400 · 0.25
Dates and quantities (temporal_numeric) 46.7 20.0
Spam detection (enron-spam) 1.000 0.920
News topic (ag-news) 0.955 0.806
Typed scores (typed-decisions) 0.438 0.395
Calibration error, product relevance 6.0 22.0
Calibration error, overall (5,096 paired) 9.5 9.3
Median latency per decision 267 ms (T4 fp16) · 280 ms (Apple M, fp32) 591 ms (hosted)

2.5× the phishing caught at McNemar p = 3e-11, and 2.1× faster with nothing leaving the machine.

Images

The hosted baseline takes no image input, so there is no like-for-like comparison. On document pages it was given Apple Vision OCR of the same images and read that text, which makes its row a two-model pipeline rather than a model that sees.

RVL-CDIP-N, 1,002 document pages, the 12 of 16 RVL-CDIP categories the set contains:

system reads accuracy ECE median latency
Apple Vision OCR + Jev 1.13.0 text 0.896 0.045 1,049 ms
Vega 0.8B, zero-shot the image 0.793 0.207 1,989 ms
DiT, best published on this set the image 0.786 n/a n/a

DiT was trained on RVL-CDIP and is the strongest published result on this out-of-distribution set. Vega has never seen the dataset and lands level with it. Not quite like-for-like: Vega chose among the 12 categories present, while a classifier trained on RVL-CDIP chooses among all 16. The pipeline latency includes the 649 ms Apple Vision takes per page, without which the API call alone is 400 ms.

MMMU-Pro, 300 items over 30 strata, chance 0.120:

system reads accuracy ECE
Gemini 3.1 Pro, best reported Oct 2026 the image 0.839 n/a
Jev 1.13.0 text only, no image input 0.337 0.160
Vega 0.8B the image 0.240 0.052
Vega 0.8B, same questions image withheld 0.157 0.062

An 800M model is nowhere near a frontier one on university-level reasoning. The result worth keeping is the last two rows: withholding the image costs 8.3 points, so the vision path carries real signal rather than the text answering alone. Calibration error is three times better than the hosted model's on the same items.

How it works

One prompt, one forward pass, then physics.

1. One formatted prompt. The state, the question and every candidate answer are written into a single sequence:

Observation: {state} Measurement ({type}): {instructions} Possible outcomes: * {option 1} * {option 2} Outcome:

The option text is genuinely part of the model's input, not metadata kept outside it. No system prompt, no chat template, no demonstrations, and the gold answer never appears. During training it exists only as a label.

2. One forward pass, five pooled vectors. The frozen language model runs once over that sequence. Vega then pools different token spans from the same hidden states: the situation span, the question span, one span per answer option, and the final token. Three questions about one state mean three prompts and three passes, except for the long-input case above, where the prefix is shared.

3. Pooled vectors become initial conditions. The situation vector projects to a world latent and gets a bounded nonlinear nudge; the question and the final token project to a probe and an impulse. Together they place a particle at position z₀ with momentum p₀ in a 64-dimensional decision space.

4. Every candidate answer becomes a valley. Each option's own features produce a Gaussian well with a centre c_k, a depth a_k and a width σ_k:

U(z) = ½κ‖z‖² − Σ_k a_k · exp( −‖z − c_k‖² / 2σ_k² )

The quadratic term keeps the particle bounded; each well pulls it toward one answer. Score levels sit on a one-dimensional rail, so ordinal neighbours are physical neighbours.

5. The particle rolls and settles. Damped Hamiltonian dynamics (symplectic Euler with friction and a learned state-space thermostat) for a fixed, small step budget, with early exit once it has settled. A decision is a short simulation with constant cost, not a sampling loop, so it is deterministic.

6. Where it settles is the answer.

E_k = ‖z_T − c_k‖² / 2σ_k² − log a_k            P = softmax(−E / τ)

What is new here

The architecture. The language model is frozen and never trained. It is perception only, read at two intermediate layers and cut off above the deepest one, so a decision never pays for the layers above it and the LM head never runs. Everything learned lives in a 55 MB engine whose state is a physical one: a position and a momentum, not a logit vector.

The training method. The engine is trained on the settling behaviour, not on next-token likelihood: a counterfactual objective pairs items that share an answer space, and a one-step world-dynamics block is trained to imagine the next state from the current one, so the latent carries what happens next rather than only what was said. Two low-rank adapters (rank 32) attach to seven engine projections, and a per-question sigmoid gate decides for each question independently whether an adapter contributes. The gate value and the chosen adapter are returned with every answer, so routing is auditable rather than implicit.

The calibration method. The readout temperature is not a constant. It is predicted per decision from the physical state the particle ended in:

log τ = b + w · [ log(1 + residual kinetic energy),
                  log(1 + distance to the nearest well bottom),
                  fraction of the step budget used,
                  log (number of options) ]

A particle still in motion, or stopped far from every well, is an uncertain decision and gets a hotter temperature. Each adapter route carries its own calibration vector. On top of that, every answer has a split-conformal set at a chosen risk level, an abstain flag below a fitted confidence floor, and an unbound flag when the particle settled far from every well, which is the model's own way of saying the question is outside what it knows.

Download counts

The repository carries a config.json so the Hub can count downloads; it counts nothing for a repository with no file it recognises. It holds no settings, nothing reads it, and the runtime still reads vega_config.json. fetch asks for it from the repository root on every load, including the 4B, which otherwise reads only from its own subfolder: a counter file that is never requested counts nothing.

Security

A checkpoint is data from a third party, so the inference code lives in this package and is reviewed with it. vegaml.load() downloads only vega_config.json, engine.safetensors and the adapter files; nothing fetched at runtime is imported or executed, and weights load through safetensors, never pickle. Tests enforce that allow-list, and CI fails on any pickle, torch.load, eval, exec, subprocess or sys.path insertion reaching the shipped package.

Repository settings worth knowing

Two things this repository cannot enforce on a private repo without a paid plan, documented here so nobody mistakes a gap for a gate:

  • Branch protection. Both classic protection and rulesets return 403 Upgrade to GitHub Pro or make this repository public, so the server enforces nothing: a force-push to main, a deletion, or a merge over a red check are all possible. The workflows still run on every push and pull request; only the enforcement is missing.

    The nearest available substitute is a local hook, .githooks/pre-push, which refuses a force-push or a deletion of main. Enable it in each clone with git config core.hooksPath .githooks. It stops the accident from a configured clone and nothing else: not another machine, not the web UI, not --no-verify. Treat it as a seatbelt, not a lock.

  • CodeQL. Code scanning needs GitHub Advanced Security on a private repository (422 Advanced security has not been purchased), so the CodeQL job reports why it skipped instead of failing forever. Secret scanning, the dependency CVE audit and the supply-chain gate do run.

Making the repository public, or upgrading, turns both on with no change to the workflows.

Licence

Apache 2.0.

Metadata

Release files for vegaml 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vegaml 0.6.0
File Size Uploaded
vegaml-0.6.0.tar.gz 61.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vegaml 0.6.0
File Interpreter ABI Platform
vegaml-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 109.5 kB

Release files / vegaml-0.6.0.tar.gz

Download URL vegaml-0.6.0.tar.gz
Size 61.1 kB
Tags Source
SHA-256 checksum
How to use checksums
439a4d487af29643c2ce4fe39f8c056b31a5235b2ae8cf84e5c708031b06cb47
BLAKE2b-256 checksum
How to use checksums
e2d82bfecee2896a29bdb98d7daba6b2b3db1744ae8bf87ce358db99fcceaefa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / vegaml-0.6.0-py3-none-any.whl

Download URL vegaml-0.6.0-py3-none-any.whl
Size 48.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e0d54bff8d34b9c567c04e6d1e6d83c8e30b4a13b927ef73ea5eca278e0838a7
BLAKE2b-256 checksum
How to use checksums
cea93a32f552c7807e66787ae375874e2bf9dc48c1a328408c6e20643f73b3e4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.1

2 release files

This release

0.6.0 This release

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page