clef-finetune
Fine-tune Cloudflare's open-source Clef and Clef-flash decision models on your own labelled decisions, then ship them as a standard Clef release folder.
Clef (Cloudflare/clef, 27B) and Clef-flash (Cloudflare/clef-flash, 9B) are Apache-2.0 decision models. You send them a state and typed questions, and they return a probability for every allowed answer in one forward pass. Cloudflare released inference code only. This repo adds the training side, following the recipe they describe in their launch post: LoRA on the backbone, the joint schema head trained alongside it, and a label-smoothed cross-entropy + Brier loss.
Status: the full pipeline is tested end to end on CPU against a tiny random model. No GPU training run on real Clef weights has been done yet, so there are no accuracy claims. See docs/KNOWN_ISSUES.md.
Features
- LoRA that actually covers Qwen3.5. Targets are found from
named_modules, so the linear-attention (Gated DeltaNet) projections get adapters, not justq_proj/v_proj. The vision tower andlm_headare excluded, and that holds when the adapter is reloaded. - Joint schema head trained alongside, in fp32 with its own learning rate, with an explicit dtype boundary to the bf16 backbone. It can also be frozen.
- Cloudflare's loss recipe: label-smoothed cross-entropy + λ·Brier for calibration, plus an optional ordinal (earth-mover's) term for score questions. The ordinal term is a supervised take on the adjacent-credit idea in RLCD, not RLCD itself.
- Schema augmentation that remaps labels: question shuffling, random question subsets, instruction paraphrases from YAML, choice option-id renaming (
approve→APPROVE/pay), and description dropout. It accounts for Cloudflare's encoder sorting choice ids, and never shuffles score levels. - Data validator: every label is checked against the legal options, with class balance reported and a warning on more than 90% majority.
- Business-grade eval: accuracy, macro-F1, NLL, Brier, ECE, and auto-decide coverage at a target precision: what share of cases the model can decide alone at, say, 99% precision, and at what confidence threshold. Reports base vs tuned as JSON + markdown.
- Release-format export:
mergewrites merged safetensors +joint_head.safetensors+ tokenizer/processor + Cloudflare's originaljoint_schema_model.py. The result loads with Cloudflare's ownload_release_model, unchanged. - Vertical starter kits: insurance claims (disposition, fraud red flags, severity, coverage line) and compliance (AML alert disposition, sanctions hit, policy violation, risk score). Both come with schemas, paraphrase templates and clearly labelled synthetic records.
- Practical training: batch size 1 + gradient accumulation, gradient checkpointing, bf16, seeded runs, checkpoint/resume, YAML configs.
- Cloudflare's code is never modified.
joint_schema_model.pyloads from your release folder or the Hub at a pinned revision. A byte-identical vendored copy (hash-checked in CI) is the offline fallback.
Install
python -m venv .venv && . .venv/bin/activate
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128 # or /whl/cpu
pip install clef-finetune
The starter kits and example configs live in the repo, so for those (or to develop):
git clone https://github.com/MersivMedia/clef-finetune && cd clef-finetune
pip install -e ".[dev]"
Needs transformers ≥ 5.10.2 (for Qwen3_5ForConditionalGeneration).
Usage
# 1. check your data (one System One request + labels per line, see docs/DATA_GUIDE.md)
clef-finetune validate data/train.jsonl data/eval.jsonl
# 2. preview augmentation
clef-finetune augment data/train.jsonl --config configs/clef-flash-insurance.yaml --copies 3 -o /tmp/aug.jsonl
# 3. measure the base model first: you may not need to train at all
clef-finetune eval --model Cloudflare/clef-flash --data data/eval.jsonl --output-dir reports/base
# 4. train (GPU), resume after interruption with --resume
clef-finetune train configs/clef-flash-insurance.yaml --output-dir runs/ins-v1
# 5. base vs tuned report
clef-finetune eval --model Cloudflare/clef-flash --data data/eval.jsonl \
--checkpoint runs/ins-v1/checkpoint-500 --compare-base --output-dir reports/ins-v1
# 6. export a release folder that Cloudflare's loader (and a System One server) can load
clef-finetune merge --model Cloudflare/clef-flash --checkpoint runs/ins-v1/checkpoint-500 --output releases/ins-v1
Try the whole chain on a laptop CPU, no downloads, using a tiny random model:
bash scripts/verify_cpu.sh # pytest + make-tiny -> validate -> augment -> train -> resume -> eval -> merge -> reload
Data format
{"id": "c1", "state": {"claim": {"description": "...", "estimate_usd": 6400}},
"questions": {"disposition": {"type": "choice", "instructions": "...", "criteria": {"approve": "...", "deny": "...", "escalate": "..."}},
"fraud_indicators": {"type": "noul", "instructions": "..."},
"severity": {"type": "score", "instructions": "...", "criteria": ["Minor", "Moderate", "Serious"]}},
"labels": {"disposition": "escalate", "fraud_indicators": true, "severity": 2}}
Config
Every setting, with defaults, is in configs/clef-flash-insurance.yaml: model path and pinned revision, augmentation probabilities, LoRA rank/alpha/target suffixes, learning rates, loss weights, checkpointing.
Docs
- Data and labelling guide: what to collect, how many examples, consensus labels, hold-out sets
- GPU sizing and RunPod how-to
- Known issues and verification status
License
Apache-2.0. Clef and Clef-flash weights and joint_schema_model.py are © Cloudflare, Inc., also Apache-2.0 (see src/clef_finetune/_vendor/NOTICE). This project is not affiliated with Cloudflare.
Metadata
Release files for clef-finetune 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| clef_finetune-0.1.0.tar.gz | 40.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| clef_finetune-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 79.5 kB
Release files / clef_finetune-0.1.0.tar.gz
| Download URL | clef_finetune-0.1.0.tar.gz |
|---|---|
| Size | 40.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9b5e3daba6d3dfd333526174dc99a454774624cfcf3fafebb912894c9cf58ba4
|
|
BLAKE2b-256 checksum How to use checksums |
0d9202243495ef18c5fa6dcf5ec2b1b1b00cf8ae845d0c66a76a4aa7ba6a0cbf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|
Release files / clef_finetune-0.1.0-py3-none-any.whl
| Download URL | clef_finetune-0.1.0-py3-none-any.whl |
|---|---|
| Size | 38.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
24f05ea95783739eb34bddd627a0b05d96478ee9e8b791deb23f0816d29fd112
|
|
BLAKE2b-256 checksum How to use checksums |
9eabe958a057eac83dcb1fa3ba2a34cc0d4dee8d94a13cbffd5a9383efedf08a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|