Updates
- [2026/09/24] Numeric fields and thinking now run batched within each batch or DAG layer; the receipt example runs 1.4–2.3x faster with thinking.
- [2026/09/24] Added image input for vision-language models, tested with Qwen3.8-27B.
- [2026/09/23] Added JevBench results: TypeLLM scored 195/231 without thinking and 228/231 with thinking.
- [2026/09/23] Added permutation averaging to improve the predictive distribution. See the blog post.
- [2026/09/22] Added
depends_ondependency graphs with incremental prefix reuse. See the blog post. - [2026/09/19] Added optional thinking mode with a per-field budget.
- [2026/09/18] Added constrained
integerandnumberoutputs.
Introduction
TypeLLM brings type-safe generation to existing autoregressive LLMs without changing their architecture or weights. Inspired by TypeSafe AI's Jev, it lets models retain their native thinking and free-form generation while producing schema-guaranteed outputs through JSON Schema. Built on SGLang, TypeLLM also supports richer interaction patterns beyond independent typed decisions.
Supported output types
String · Integer · Number · Boolean · Enum choice — See schemas and examples.
Features
- No out-of-schema hallucinations — Choices stay within the allowed values.
- Negligible output-token cost — Single-token categorical selection and bounded numeric decoding; optional thinking adds tokens.
- Shared-prefix reuse — KV caching avoids reprocessing shared context.
- Dependency-aware execution — Run decisions sequentially, batch independent fields, or declare
depends_onto form a dependency graph. - Made for open autoregressive LLMs — Use compatible models you already serve with SGLang.
- Supports thinking mode — Enable reasoning before the final constrained answer.
- Image input — Pass images to vision-language models alongside the text context. See Image input.
- Permutation averaging — Reduce option-order bias on explicit enum questions with sampled or exhaustive orderings. See the docs.
JevBench results
Evaluated on 231 public JevBench tasks.
Full results and all per-task answers · Method and configuration
Quick start
1. Serve a model with SGLang
Use SGLang to configure and serve a compatible autoregressive model on your local GPU server. This example uses Qwen3.8-27B; follow the Qwen3.8-27B SGLang deployment guide to start it with prefix caching enabled.
See Supported models for tested checkpoints and thinking behavior.
2. Run TypeLLM
pip install -U typellm
Point TypeLLMClient at the SGLang server's HTTP endpoint:
from typellm import TypeLLMClient
client = TypeLLMClient(
"http://127.0.0.1:30000",
model="Qwen/Qwen3.8-27B",
)
Example request:
result = client.generate(
context="""
Receipt from Hilton London
Total: £324.50
Employee travelled to London for a client meeting.
""",
questions={
"merchant": {
"type": "string",
"instructions": "Return only the merchant name.",
},
"total": {
"type": "number",
"instructions": "Extract the total amount in GBP.",
},
"expense_type": {
"type": "string",
"enum": ["meal", "travel", "equipment"],
"instructions": "What type of expense is this?",
},
"reimbursable": {
"type": "boolean",
"instructions": "Should this expense be reimbursed?",
},
"confidence": {
"type": "number",
"enum": [0.0, 0.25, 0.5, 0.75, 1.0],
"instructions": "How confident are you?",
},
},
)
print(result)
Example return:
{
"merchant": "Hilton London",
"total": 324.5,
"expense_type": "travel",
"reimbursable": True,
"confidence": 0.75,
}
Output types
TypeLLM supports finite decisions, numeric fields, and free text:
| Field | Schema | Returned value |
|---|---|---|
| Text | {"type": "string"} |
str |
| Integer | {"type": "integer"} |
int |
| Number | {"type": "number"} |
float |
| Boolean | {"type": "boolean"} |
bool |
| Enum choice | {"type": "string", "enum": ["meal", "travel"]} |
Candidate type: str, int, or float |
Enum choices support string, integer, and number types, with at most 24 values. The declared type validates the candidate values.
A string without enum generates free text:
result = client.generate(
context="The train ticket is for a client meeting.",
questions={
"summary": {"type": "string", "instructions": "Summarize in one sentence."},
},
)
Ask for a numeric answer without enumerating every possible value:
result = client.generate(
context="Calculate the requested value accurately.",
questions={
"answer": {
"type": "number",
"instructions": "What is 17.5 multiplied by 4?",
},
},
)
print(result)
# {"answer": 70.0}
These fields generate plain
decimal notation with at most 32 digits by default; set
TypeLLMClient(numeric_max_digits=...) to adjust this limit.
Use instructions to tell the model what decision to make:
{
"type": "string",
"enum": ["billing", "technical", "account"],
"instructions": "Which team should handle this ticket?",
}
If instructions is omitted, TypeLLM uses description or an instruction
generated from the field name.
Thinking mode
Thinking is off by default. Enable it when constructing the client:
client = TypeLLMClient(
"http://127.0.0.1:30000",
model="Qwen/Qwen3.8-27B",
thinking=True,
)
result = client.generate(context=context, questions=questions)
Set thinking_budget=2048 to cap reasoning per field; no explicit budget is set
by default. If reasoning reaches its limit or ends early at a recognized turn
terminator, TypeLLM closes a nonempty thinking block and proceeds to the typed
answer. Empty unfinished reasoning and unrecognized stops raise an error.
Models with always-on thinking still reason with thinking=False;
thinking_budget applies to them too. See Supported models.
Image input
Pass images with images= alongside the text context. The served model must be
a vision-language model, such as Qwen/Qwen3.8-27B.
result = client.generate(
context="The customer says this receipt was charged twice.",
images=["receipt.png"],
questions={
"total": {"type": "number", "instructions": "What is the receipt total?"},
"paid": {"type": "boolean", "instructions": "Is the receipt marked as paid?"},
},
)
Each image can be a local file path, an http(s) URL, a data: URI, raw bytes,
or a PIL image. Local files are read by the client, so the SGLang server does
not need access to your filesystem. Images come before the text in the first
user turn, and every request in the call carries them, across batch,
sequential and DAG execution, permutations and thinking.
TypeLLM reads the image placeholder from the model's chat template. If the
template does not render image content, generate() raises an error before
sending any request.
Dependency-aware execution
The default execution="auto" selects batch execution unless a field declares
depends_on.
| Mode | Field context | Execution order |
|---|---|---|
batch |
Original context only | Independent fields run together |
sequential |
Original context and all earlier answers | Field declaration order |
dag |
Original context and dependency results | Dependency order |
Set the mode per request or on the client:
result = client.generate(
context=context,
questions=questions,
execution="sequential", # Or "batch" for independent fields
)
Batch execution shares the cached context across branches. Configure SGLang's
--max-running-requests for the desired concurrency.
Dependency execution (depends_on)
Declare which earlier results a field needs. Forward references are allowed: fields do not need to be declared in execution order.
result = client.generate(
context="The payments service is returning errors after a deployment.",
questions={
"system": {
"type": "string",
"enum": ["payments", "accounts", "search"],
"instructions": "Which system is affected?",
},
"severity": {
"type": "string",
"enum": ["low", "medium", "high"],
"instructions": "Assess severity for the affected system.",
"depends_on": ["system"],
},
"deployment_related": {
"type": "boolean",
"instructions": "Is the incident related to a deployment?",
"depends_on": ["system"],
},
"rollback": {
"type": "boolean",
"instructions": "Based on the incident assessments, should we roll back?",
"depends_on": ["severity", "deployment_related"],
},
},
)
This runs system, then severity and deployment_related, then rollback.
Each layer finishes before the next starts; unrelated branches remain separate.
depends_onlists unique field names. Missing or empty lists mark independent roots.- Fields receive their direct and transitive dependency results. Probability-returning dependencies contribute only their selected value.
- Unknown names, self-dependencies, duplicates, and cycles raise
SchemaError. - Any
depends_on, including[], activates DAG execution inautomode. Combining it with explicitbatchorsequentialraisesSchemaError. - Both
questionsand object-formschema.propertiessupport dependencies. Returned keys follow field declaration order.
All fields execute. Dependencies do not change enums, substitute values into instructions, or conditionally skip fields. A failed layer stops subsequent layers.
Incremental prefix reuse along dependencies
TypeLLM extends parent prompts along dependency paths, retaining previous answers
and reasoning for KV cache reuse. In a chain A → B → C, each step builds on the
previous prefix; independent branches share their common prefix.
When a field depends on multiple parents, TypeLLM reuses one parent prefix and includes all dependency results. SGLang manages the cache; KV tensors from different branches are not merged.
Batch performance
A local run with Qwen3.8-27B NVFP4 on one NVIDIA RTX PRO 6000 Blackwell GPU used
roughly 1,100 context tokens and 16 Boolean fields, with
--max-running-requests 16.
| Execution | End-to-end latency | Latency per decision | Relative throughput |
|---|---|---|---|
| Sequential | 9.35 s | 0.584 s | 1.0x |
| Batch | 1.61 s | 0.101 s | 5.8x |
Each branch reused 1,088 cached tokens. Within a batch or a DAG layer, strings and choices are each generated in one batched request, numeric fields decode in lockstep with one batched request per digit, and thinking runs for every field in one batched request. Results depend on the model, workload, and server configuration. Sequential fields see earlier answers; batch fields are independent, so the modes serve different workflows.
Probabilities and sampling
Set return_probabilities on individual enum or boolean fields:
result = client.generate(
context=context,
questions={
"expense_type": {
"type": "string",
"enum": ["meal", "travel", "equipment"],
"return_probabilities": True,
},
},
)
{
"expense_type": {
"value": "travel",
"probabilities": {
"meal": 0.04,
"travel": 0.93,
"equipment": 0.03,
},
}
}
Only opted-in fields return value and probabilities; other fields return plain values.
The option is not supported on open Numeric or Text fields.
Per-question permutation averaging
Add permutations to an enum question to reduce option-order bias. TypeLLM averages the probabilities and keeps the same return format.
result = client.generate(
context="A single roll of a fair die.",
questions={"roll": {
"type": "string",
"enum": ["one", "two", "three", "four", "five", "six"],
"instructions": "What number will come up on this roll?",
"permutations": 8,
"return_probabilities": True,
}},
)
Use 8 for eight distinct orderings or "all" for every ordering (up to 720). Omit it or use 1 to keep the original behavior. Only explicit enum fields support this option.
Argmax is the default. To enable sampling:
client = TypeLLMClient(
"http://127.0.0.1:30000",
mode="sample",
temperature=0.8,
seed=42,
)
Sampling applies to finite candidates for Choice fields and to token generation
for Numeric and Text fields. temperature controls sampling in each case.
For a one-off request, use the convenience function:
from typellm import run_schema
result = run_schema(
context=context,
questions=questions,
base_url="http://127.0.0.1:30000",
model="Qwen/Qwen3.8-27B",
)
Cost analysis
Enum and boolean fields use one output token each. Numeric, text, and optional thinking outputs use multiple tokens.
For a sequential workflow with D fields, C original context tokens, and
roughly S new tokens per turn, input prefill counts are:
without prefix reuse: O(D*C + D^2*S)
with prefix reuse: O(C + D*S)
For independent batch fields, the shared context is prefilled once, followed by each field's question. These counts describe input token positions, not GPU compute or latency: new tokens still attend to the cached prefix, and reuse depends on cache availability. Hosted billing depends on the provider's cached-input pricing.
Supported models
The following models have been tested with TypeLLM on a live SGLang GPU server.
| Model / checkpoint | Thinking support |
|---|---|
Qwen/Qwen3.8-27B |
On / off |
Qwen/Qwen3.5-0.8B/4B/9B |
On / off |
openbmb/MiniCPM5-1B |
On / off |
inclusionAI/Ling-mini-2.0 |
Off only |
inclusionAI/Ring-mini-2.0 |
Always on |
Other sizes in the Qwen3.5 and Qwen3.8 families are expected to be compatible.
Image input has been tested with Qwen/Qwen3.8-27B.
The MiniCPM5, Ling and Ring runs used an RTX PRO 6000 Blackwell and SGLang 0.5.19 on 2026-09-22.
Use the checkpoint ID as model=. If the server's tokenizer path is unavailable
locally, set tokenizer= to its matching Hugging Face ID or local directory.
The tokenizer must load from standard artifacts without custom model code.
Comparison with Jev-style models
| Feature | TypeLLM | Jev | openjev-sglang | system-one-open | OpenJev DeBERTa |
|---|---|---|---|---|---|
| Enum selection | ✓ | ✓ | ✓ | ✓ | ✓ |
| Boolean decisions | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rubric scoring | Numeric enum; no dedicated Score API | Score | Score | Score | Score |
| integer/decimal type | ✓ | — | — | — | — |
| string type | ✓ | — | — | — | — |
| Enable Thinking | ✓ | — | — | — | — |
| Image input | ✓ | Not documented | — | Not documented | — |
| Multi-field execution | Batch, sequential, DAG | Batch | Batch | Batch | Batch |
| Built-in field dependency graph | ✓ | — | — | — | — |
| KV prefix reuse | Shared context + dependency paths | Not disclosed | Shared context | Not documented | Not applicable |
© 2026 TypeLLM
Release files for typellm 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| typellm-0.1.5.tar.gz | 61.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| typellm-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 103.0 kB
Release files / typellm-0.1.5.tar.gz
| Download URL | typellm-0.1.5.tar.gz |
|---|---|
| Size | 61.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7d9ebb637f4961b5f84b2bf238012bec512843212a005452bbe82840e9554c73
|
|
BLAKE2b-256 checksum How to use checksums |
9a39a3bdd218d429de21282aacfc15ab4542a9cffc66ccb46f68688ddb2015f3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency logRelease files / typellm-0.1.5-py3-none-any.whl
| Download URL | typellm-0.1.5-py3-none-any.whl |
|---|---|
| Size | 41.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7c00a1e844ade182464376f8e1cbb2e829d5ee73d5c519dd40926ed70e43e3ac
|
|
BLAKE2b-256 checksum How to use checksums |
6a93a161e1ef855591df6c8a5515c24bc72813647cd2a5451169af43493fd654
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency log