sokudan (即断)
A Japanese System One decision model: send text and typed questions, get typed answers with probabilities, without generating a single token.
日本語: README_ja.md
Why
- No generation. Answers are read off a decision head's logits. There is no text to parse, no JSON to validate, and no way to answer outside the options you sent.
- Japanese-native. The backbone is
sbintuitions/modernbert-ja-310m, and the model is trained and evaluated on Japanese. On Japanese business messages it beatslaya-multilingualon every metric below. - One forward pass per question. No decoding loop and no retries. 314.6M parameters, so it also runs on a CPU.
Quickstart (30 seconds)
Python 3.11.
pip install sokudan
Development version (the main branch): pip install git+https://github.com/hiroki-abe-58/sokudan.git.
What pip install sokudan brings depends on the platform (v0.3.0):
| platform | array library installed | sokudan.load() runs on |
|---|---|---|
| Apple silicon, macOS 14 or later | MLX (mlx>=0.32.2,<0.33); no torch |
MLX, float16 |
| Linux, Windows, Intel Mac, Apple silicon on macOS 13 | torch | torch: cuda, then mps, then cpu |
pip install "sokudan[torch]"adds torch on any platform (forbackend="torch"on Apple silicon, and for the training and evaluation scripts).- To use a GPU on Windows / Linux, install a CUDA build of torch first (for example
pip install torch --index-url https://download.pytorch.org/whl/cu128), then install sokudan. pip install "sokudan[mlx]"names MLX explicitly (same platforms as above).pip install "sokudan[serve]"adds the server.
import sokudan
agent = sokudan.load("GeneLab/sokudan-ja-310m")
questions = {"department": {"type": "choice", "instructions": "この問い合わせはどの部署が担当すべきか",
"criteria": {"請求": "支払い・返金", "技術": "不具合・障害", "営業": "料金・新規契約", "その他": "上記以外"}}}
result = agent.predict("先月の請求で同じ金額が二回引き落とされています。至急ご確認ください。", questions)
print(result["answers"]["department"]["choice"])
Expected output:
請求
result["answers"]["department"]["probabilities"] holds the full distribution over the four options.
Question types are choice, score (an ordinal scale) and noul (P(yes); bool is an alias). The schema is free per request; no retraining.
Pass the state as a string. A dict is rendered as key: value lines, which is not the input the model was trained on, and the output changes.
Backends (v0.3.0). sokudan.load(..., backend="auto" | "mlx" | "torch", dtype=None | "float16" | "float32"). auto tries MLX (Apple silicon with mlx installed), then torch on mps, cuda and cpu; each candidate answers one short self-check request, and a failure is a warning followed by the next candidate. An explicit backend or device does not fall back. agent.backend says which one is in use. MLX runs float16 by default (the heads stay float32); on bench_ja, MLX float32 and float16 give the same four metrics as torch to three decimals. Details and measurements: docs/mlx.md.
Calibration (since v0.2.1). load applies one temperature to noul/bool answers by default (the calibration.json shipped with the weights); choice and score probabilities are raw. sokudan.load(..., temperatures=None) turns it off. Each result says which answers were calibrated (calibrated, calibrated_answers).
The same steps in a notebook on a free CPU runtime: the three question types, calibration on and off, and
sokudan serve called with curl (notebooks/sokudan_quickstart.ipynb).
bench_ja: v0.2 against laya-multilingual
300 Japanese business messages, three unseen schemas (4-way department routing, 3-level urgency, churn suggestion). Uncalibrated. Each model row is a single run.
| choice acc | score RPS↓ | bool acc | bool AUROC | |
|---|---|---|---|---|
| sokudan-ja-310m v0.2 | 0.880 | 0.075 | 0.780 | 0.844 |
laya-multilingual (ja) |
0.747 | 0.232 | 0.543 | 0.523 |
| majority class | 0.380 | 0.197 | 0.703 | — |
| random | 0.253 | 0.201 | 0.513 | — |
bench_jaandbench_enship in this repo (data/bench_ja.jsonl,data/bench_en.jsonl) under CC BY 4.0, separate from the code's Apache-2.0. Please use them for evaluation, not training (a request, not a licence restriction).- v0.2's score accuracy is 0.817. Every metric, v0.1's three-seed figures and
bench_enare in the model card anddocs/benchmarks.md.
How v0.2 was made
- A model soup of eight seeds. Seeds 0-7, each trained exactly as v0.1 was, averaged tensor by tensor in float32 (heads included). Architecture, data and inference code are v0.1's; inference costs one model.
- Chosen by a rule fixed in advance. Four candidate soups; the rule (largest held-out M1m, no guardrail regressions) was committed before any
bench_jarun (docs/release_candidate.md,docs/research_protocol.md). bench_jawas measured once, after the release rule was committed: bool AUROC +0.01 or better, score RPS −0.005 or better, and choice / bool / score accuracy within −0.01 of v0.1. All were met; bool accuracy is the one that went down (−0.008).- Reproducing it.
scripts/make_soup.pyrebuilds the soup from the eight checkpoints and checks every SHA-256. The publishedmodel.safetensorsis8750a833…5b965(full hashes in the model card). The member checkpoints are not distributed, and retraining does not reproduce them bit for bit. - v0.1 remains available:
sokudan.load("GeneLab/sokudan-ja-310m@v0.1").
Limits
- Position sensitivity. On a four-level
score(condition E of the position probe) v0.2 chose the first option for 5 of 300 items, accuracy 0.347. Do not extrapolate from three levels to four or more. On held-out states the first slot is still slightly disfavoured: first-slot rate over all orders 0.239 (score) and 0.289 (choice), against 1/3 if order did not matter. Three kinds of fix were tried (averaging over orders at inference, a permutation-KL term in training, shuffling option order in the training data); each moved some position checks but none kept held-out accuracy non-inferior, so none shipped (model card, Limits). A catch-all option ("その他", "Other") is chosen less when it sits first or last; ordinary options hardly depend on position onbench_ja(all 24 orders of the four departments: each slot chosen 0.243–0.256 of the time). - Catch-all options are under-chosen. The argmax rarely picks "その他 / Other" even when it is right:
- On
bench_ja, "その他" recall is 0.289 (11 of 38;bench_en0.354). Misses go mostly to 技術 / Technical, while precision is 0.846. - In a small probe (90 expense descriptions by one author, 10 accounting categories + "その他: 上記以外"), it was chosen for 0.295 of the states that fit no category with random option orders, and 0.155 with "その他" last.
- P(その他) still ranks those states well (AUROC 0.94). Read P(catch-all) and set your own threshold; do not put the catch-all first or last. Guide:
docs/choice_guidance.md. - Near-miss pairs that people also split (消耗品費 / 事務用品費; 会議費 / 交際費 for meals with clients) are not separated either.
- On
boolunder-predicts true. Mean P(true) 0.133 against a gold rate of 0.297 (0.198 with the default bool calibration). The ranking works (AUROC 0.844); set the threshold from your own prior. Calibration does not move the 0.5 threshold.boolis calibrated by default (one temperature, ECE 0.181 → 0.105 onbench_ja);scoreandchoiceare left raw because the score temperature fitted on validation made bench RPS worse (0.075 → 0.132). Details:docs/calibration.md.choicewas re-checked for v0.3.0 and stays raw: the validation temperature loweredchoiceECE onbench_ja(0.088 → 0.066) but raised it onbench_en(0.091 → 0.228).- Latency grows with the number of questions. The state is re-encoded for every question.
- Long states lose accuracy. The backbone's local attention window is 128 tokens; training states average 134 tokens (p95 237).
- Japanese only. On
bench_en, bool accuracy is 0.690, level with the majority class (0.683). - Synthetic data. Training data (21 domains, 4,833 documents, 31,243 labelled pairs) and
bench_jaare both generated by one LLM. Not measured on other tasks. Do not use it to replace a human decision about hiring, credit, discipline, medicine or law.
The full list is in README_ja.md and the model card.
/v1/systemone-compatible server
pip install "sokudan[serve]"
sokudan serve --port 8000
Development version: pip install "sokudan[serve] @ git+https://github.com/hiroki-abe-58/sokudan.git".
POST /v1/systemone takes the same request and returns the same answer shapes as TypeSafe's public API reference, so a client written for that format can point its base URL at http://127.0.0.1:8000. GET /health reports the loaded model, calibration, and how a JSON state is rendered.
Built from public documentation and the examples in open implementations' READMEs; not affiliated with or endorsed by TypeSafe AI, and this repository never calls their service.
Guide: docs/serving.md. Field-by-field table: docs/systemone_wire_format.md.
Used by
- Expense account suggestion in an accounting app (PoC): given a purchase description, sokudan ranks 1–3 candidate accounts for a human to confirm, running on an M1 Max. Thread on X
More
- README_ja.md: the Japanese README, with the design notes (joint encoding, the dynamic-K cumulative link) and every limit.
docs/architecture.md,docs/baseline_ja.md,docs/baseline_lev.md(against lev on the same machine and harness),CHANGELOG.md.- Development:
uv sync --extra dev,uv run pytest,uv run ruff check .. From v0.3.0 the dependencies only data building, training and evaluation use (datasets,fugashi,unidic-lite,matplotlib) are in thetrainextra;devincludessokudan[train], souv sync --extra devstill installs them (uv sync --extra train/pip install "sokudan[train]"without the dev tools). On Apple silicon with macOS 14+, add--extra torchfor the tests and training that need torch.
Apache-2.0. The backbone sbintuitions/modernbert-ja-310m is MIT (see NOTICE).
Release files for sokudan 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sokudan-0.3.0.tar.gz | 249.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sokudan-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 424.9 kB
Release files / sokudan-0.3.0.tar.gz
| Download URL | sokudan-0.3.0.tar.gz |
|---|---|
| Size | 249.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
15c3e67b56a3aebf3496ecf896a34dfc6df2c98eb1658380f794d3b074f13c0d
|
|
BLAKE2b-256 checksum How to use checksums |
e4881a5b5b9aa4ca08a43fde22016f5d179ccc3e0c73f824230f567af3f06a44
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|
Release files / sokudan-0.3.0-py3-none-any.whl
| Download URL | sokudan-0.3.0-py3-none-any.whl |
|---|---|
| Size | 175.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8873a9f14cf2ad3cd587dafa76ea73bc63a139add955374375d6d7085a70c5b3
|
|
BLAKE2b-256 checksum How to use checksums |
7fcd06bf1fc5e5db5a0aa104191ea0f50470e0688953168606e85770461293ac
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|