Skip to main content

taskdistill

Replace an expensive LLM API call on a narrow task with a small fine-tuned model on your Mac, and prove it with numbers. taskdistill captures the traffic your application already sends, curates it into training data, LoRA fine-tunes a 0.5B–1.5B Qwen2.5 student with MLX on Apple Silicon, evaluates it against the original model on held-out data, and serves an OpenAI-compatible cascade: the student answers first and hands the request to the original model (the teacher) when its confidence is below a threshold chosen on validation data. Your application changes only its base URL.

taskdistill demo banking77, then a request to the cascade server

Why

Many production LLM calls are narrow: route a support message to one of N intents, pull eight fields out of an invoice e-mail. They ship on a mid-tier API model because that is the fastest way to launch. At volume the call becomes a steady cost and a latency floor (in this project the teacher answered in about 594 ms at p50 and 1,242 ms at p95 (Banking77 labelling calls, 2026-09-26)), and it sends every input to a third party. A small fine-tuned model could take most of that traffic, but teams hold back for three reasons: nobody has the logged data in trainable shape, nobody trusts the small model's quality, and nobody knows which requests it will get wrong.

taskdistill is for the engineers who own such a call. It gives you a reproducible way to

  1. collect training data from the traffic you already have (a capture proxy, or an import of existing logs),
  2. measure the student against the teacher on held-out data, with baselines and confidence intervals, and
  3. run a cascade whose quality/cost trade-off you pick explicitly, on validation data, instead of hoping for it.

How it relates to other tools

Project What it does What it does not do here
OpenPipe Hosted fine-tuning of smaller models from logged requests; acquired by CoreWeave in 2025, its platform stopped new training and inference on 30 July 2026 and moved to Weights & Biases. Not local; its open-source repository has been paused since 2024.
Predibase / LoRAX Managed fine-tuning and serving (Predibase, acquired by Rubrik in 2025); LoRAX serves many LoRA adapters on one NVIDIA GPU. LoRAX serves only; it needs an NVIDIA GPU on Linux and does not capture, train or evaluate.
RouteLLM Routers that send each query to a strong or a weak existing general-purpose model. It does not train a task-specific student from your traffic.
FrugalGPT Research method and code for a cascade over a sequence of paid LLM APIs. It cascades between existing API models; no local student, no capture proxy.
distilabel Pipelines that generate synthetic data and AI feedback and emit datasets. It neither trains nor serves models.
LiteLLM Gateway exposing 100+ LLM APIs in the OpenAI format, with routing, cost tracking and logging; fine-tuning passes through to hosted APIs. It does not build training sets from its logs or distil locally.

The gap taskdistill fills: one local, open-source pipeline from captured traffic to a cascade with a validation-chosen threshold on Apple Silicon, with an evaluation report you can reproduce offline. (Statements checked against each project's own site or repository on 2026-09-26.)

Quickstart

On an Apple Silicon Mac (macOS 14 or later) with uv installed. No API key is needed: without one, the demo replays the recorded teacher outputs that ship with the package.

uvx --from git+https://github.com/B0yko/taskdistill taskdistill demo banking77

Measured quick-profile demos in replay mode, each in a fresh workspace and directory (2026-09-27; reports/demo_timing.json):

Machine Demo Cache Install Demo Total Load average before
MacBook Air, Apple M5, 24 GB banking77 warm 12.9 s 2.4 min 2.7 min 1.90/1.98/2.02
MacBook Air, Apple M5, 24 GB invoices warm 0.2 s 2.1 min 2.1 min 2.62/2.48/2.23
Mac Studio, Apple M4 Max, 128 GB banking77 warm 3.1 s 1.4 min 1.5 min 7.43/7.80/8.05
Mac Studio, Apple M4 Max, 128 GB invoices warm 0.2 s 1.1 min 1.1 min 12.43/10.15/8.98
Mac Studio, Apple M4 Max, 128 GB banking77 cold 7.1 s 1.6 min 1.7 min 9.31/9.69/8.89

The cold run (empty uv cache and Hugging Face cache on the Mac Studio, Apple M4 Max, 128 GB) took 1.7 min: 7.1 s to install and 1.6 min for the demo, including the 0.29 GB base-model download. Every measured run finished in under five minutes; slower Macs and slower connections will take longer.

The demo runs the whole pipeline on Banking77 and leaves a trained student, a report in reports/banking77/ and a request.json in the current directory. Serve the cascade and call it:

uvx --from git+https://github.com/B0yko/taskdistill taskdistill serve --task banking77
curl -s -D - http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d @request.json

The response is a standard chat.completion; the x-taskdistill-route header says whether the student or the teacher answered and x-taskdistill-confidence carries the student's confidence. taskdistill demo invoices does the same for JSON extraction on synthetic invoices. Both demos use the quick profile, sized to finish in a few minutes: Banking77 trains on 2,000 examples (300 validation, 500 test, 200 iterations) and invoices on 120 documents (24 validation, 36 test, 6 per layout, 30 iterations), cut down from 400/60/100 so the demo stays within the time budget. Use --profile full for the configuration behind the numbers below (it takes tens of minutes to hours on a laptop).

How it works

flowchart LR
  subgraph today[Your application today]
    A[App] -->|base URL| P[capture proxy]
    P --> T[(teacher API)]
  end
  P --> S[(SQLite store)]
  S --> C[curate] --> D[data/*.jsonl] --> TR[train: LoRA with MLX] --> AD[adapter]
  AD --> E[eval]
  E -->|validation: run, threshold, calibration| TH[selected_run.json + threshold.json]
  E -->|test: report only| R[report]
  subgraph cascade[Your application after]
    A2[App] -->|base URL| CS[cascade server]
    CS -->|confidence >= t| ST[student on the Mac]
    CS -->|confidence < t| T2[(teacher API)]
  end
  TH --> CS
  CS --> S
  S --> R
Step Command What it does
Capture taskdistill capture --task T OpenAI-compatible reverse proxy: forwards POST /chat/completions unchanged with the client's own Authorization, and stores request and response bodies, usage, latency and status. It never stores a header. --import loads existing logs instead.
Curate taskdistill curate --task T Extracts the task input, merges records by input hash (this is how gold labels reach captured inputs), normalises teacher outputs, scrubs PII, assigns splits, removes duplicates and near-duplicates inside and across splits, labels unlabelled inputs with the teacher, drops (never truncates) over-long examples, and fails unless an independent leakage check finds zero duplicates across splits.
Train taskdistill train --task T LoRA on the 4-bit base with mlx-lm, loss on completion tokens only, keeping the checkpoint with the best validation loss. --backend torch uses transformers + PEFT instead.
Eval taskdistill eval --task T --select Chooses the run and the cascade threshold on validation, then scores student, teacher, cascade and baselines on test with paired bootstrap confidence intervals and calibration metrics.
Serve taskdistill serve --task T FastAPI server with /v1/chat/completions, /v1/models, /healthz and /metrics. Low-confidence, unsupported and unparsable requests go to the teacher with the teacher's key, never the client's.
Report taskdistill report --task T Quality, operating point, $/1k requests, latency and break-even volume, optionally from real served traffic (--from-serve-log).

Confidence. Classification decodes greedily under a trie of the canonical labels, so every answer is a valid label; confidence is the product of the renormalised probabilities of the chosen tokens (the end token included, so a label that is a prefix of another is handled). Extraction generates JSON freely; each field's confidence is the product of the probabilities of the tokens spanning its value, and the document's confidence is the minimum over fields (0 for invalid JSON or a schema violation).

Selection on validation only. Every choice (base model, seed, checkpoint, threshold, isotonic calibration, baseline hyperparameters) is made by a function that accepts only a ValidationSplit; passing a TestSplit raises TypeError. The threshold is the one with the lowest escalation rate that meets the target on validation, and the test split then reports whether the target held.

Use it on your own call

Say your application classifies support tickets with a paid API model.

  1. Install the command and create a task spec.

    uv tool install git+https://github.com/B0yko/taskdistill
    
    taskdistill init tickets --type classification
    

    Edit tasks/tickets/teacher_prompt.md (the system prompt your application sends today), labels.txt, and in task.yaml the teacher's model, base_url and, for OpenRouter, the pinned provider in extra_body. Everything below runs in the directory that contains tasks/; the workspace defaults to ./.taskdistill.

  2. Capture traffic. Start the proxy and change only the base URL of your OpenAI client:

    export TASKDISTILL_TEACHER_BASE_URL=https://openrouter.ai/api/v1   # where the proxy forwards
    taskdistill capture --task tickets
    
    from openai import OpenAI
    
    client = OpenAI(base_url="http://127.0.0.1:8787/t/tickets/v1", api_key=YOUR_KEY)  # the key is passed through
    

    Or import logs you already have (--format openai for request/response pairs, pairs for input/output, inputs for unlabelled inputs that the teacher will label). Gold labels, if you have any, are imported with --format inputs and joined to captured traffic by input hash:

    taskdistill capture --task tickets --import gold.jsonl --format inputs
    
  3. Curate, train and evaluate. Curate prints a projected cost from a 50-request sample before labelling anything and asks for --yes above $0.50; --max-usd caps the run.

    export TASKDISTILL_TEACHER_API_KEY=...    # or OPENROUTER_API_KEY
    taskdistill curate --task tickets --max-usd 2
    taskdistill train --task tickets
    taskdistill eval --task tickets --select
    

    Without gold labels, eval reports agreement with the teacher, and the default cascade target (97% agreement) needs no gold at all.

  4. Serve the cascade and point your client at it. The model string your application sends can stay as it is.

    taskdistill serve --task tickets
    
    client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
    
  5. Watch it. taskdistill report --task tickets --from-serve-log --since 24h recomputes the figures from real served traffic, including the observed escalation rate; a rate far above the validation estimate means the traffic has drifted away from the training data.

Results

Numbers below come from scripts/reproduce.sh (full profile, recorded teacher outputs) on 2026-09-27 on a MacBook Air, Apple M5, 24 GB (Mac17,4), macOS 26.6.2, plus one extra zero-shot evaluation (the invoices table also shows the 1.5B zero-shot base); teacher outputs were recorded on 2026-09-26. The test splits were scored 8 times (Banking77) and 9 times (invoices), each time by an evaluation shown in these tables; no choice used them: every choice (base model, seed, threshold, isotonic calibration) was made on the validation split alone. Figures marked “recorded” below (the teacher and cascade rows) replay the teacher outputs captured then, not a live call. The live bench and latency numbers are dated measurements and are not expected to reproduce exactly on different hardware or under different load. Unless noted otherwise, every score has a 95% paired bootstrap interval in brackets (1,000 resamples).

Two machines were used. Quality, calibration and training numbers come from the MacBook Air, Apple M5, 24 GB that trained the students; it is fanless and shared with other work. Cost, latency, the live bench and the cross-machine reproduction come from a Mac Studio, Apple M4 Max, 128 GB, used only for those runs. Each table names its machine.

Bottom line. Each cascade's threshold was chosen on validation; on test:

  • Banking77. Target held on test: yes — Agreement against the teacher was 97.6% on test (target 97.0%) at a 23.7% escalation rate, close to the 97.1% on validation at 23.4%.
  • Invoices. Target held on test: no — Agreement against the teacher met the 97.0% target on validation (98.5% at 0.0% escalation) but fell to 94.1% on test (1.7% escalation).

Metrics: agreement is how often a system gives the teacher's answer (for invoices, field micro-F1 against the teacher's JSON); accuracy and macro-F1 (F1 averaged over the 77 intents, each counted equally) are against the gold labels; ECE (expected calibration error, lower is better) is the average gap between the student's stated confidence and how often it is right; AUROC (0.5 = chance, 1.0 = perfect) is how well that confidence separates right answers from wrong ones. For invoices, JSON validity is the share of outputs that parse as a JSON object, field micro-F1 and field EM (exact match) score the 8 fields against the gold, and Doc EM is the share of documents with all 8 fields right.

Banking77 (77 intents)

Test split, n = 3,075. Cells are the score and its 95% paired bootstrap interval; a row aggregating more than one seed shows mean ± sample standard deviation only (see the per-seed table below for each seed's own score, and interval where the eval recorded one). ECE and AUROC use raw (pre-isotonic) student confidence.

System n Accuracy Macro-F1 Agreement ECE AUROC
teacher (deepseek/deepseek-v4.1-flash) 3,075 75.8% [74.2, 77.4] 74.9% [73.0, 76.1] reference — —
base 0.5B zero-shot (labels in the prompt) 3,075 23.0% [21.6, 24.5] 20.5% [19.0, 21.6] 25.6% [24.1, 27.0] 35.3% [33.9, 36.7] 0.738 [0.718, 0.761]
TF-IDF + logistic regression (teacher labels) 3,075 72.5% [70.8, 74.0] 71.4% [69.7, 72.5] 83.5% [82.2, 84.8] — —
student qwen2.5-0.5b (teacher labels, 3 seeds) 3,075 75.7% ± 0.5 74.6% ± 0.7 88.1% ± 0.2 15.2% ± 0.3 0.801 ± 0.003
student qwen2.5-1.5b (teacher labels) 3,075 76.5% [74.9, 78.0] 75.5% [73.9, 76.7] 89.2% [88.1, 90.4] 15.4% [14.1, 16.9] 0.797 [0.779, 0.814]
student qwen2.5-0.5b (gold labels) 3,075 92.0% [91.0, 92.9] 92.0% [91.0, 92.9] 75.8% [74.2, 77.4] 1.4% [1.0, 2.4] 0.915 [0.897, 0.931]
cascade (qwen2.5-0.5b-full-s13, t = 0.874) 3,075 76.1% [74.5, 77.7] 75.0% [73.3, 76.2] 97.6% [97.0, 98.1] — —

Per-seed scores for student qwen2.5-0.5b (teacher labels, 3 seeds) (selected run: qwen2.5-0.5b-full-s13):

Run Seed Accuracy Macro-F1 Agreement ECE AUROC
qwen2.5-0.5b-full-s13 (validation-selected) 13 75.5% [73.8, 77.0] 74.4% [72.7, 75.6] 88.1% [86.9, 89.3] 15.0% [13.7, 16.5] 0.799 [0.781, 0.816]
qwen2.5-0.5b-full-s14 14 75.3% 74.1% 87.9% 15.6% 0.801
qwen2.5-0.5b-full-s15 15 76.3% 75.4% 88.2% 15.0% 0.804
Value
Target Agreement ≥ 97.0% against the teacher
Chosen threshold (on validation) 0.8740
Escalation rate, valid / test 23.4% / 23.7%
Cascade − teacher, Agreement against the teacher (target) -2.4 pts [-3.0, -1.9]
Cascade − teacher, Accuracy (gold) +0.3 pts [-0.1, +0.8]
Target held on test yes

Target held on test: yes — Agreement against the teacher was 97.6% on test (target 97.0%) at a 23.7% escalation rate, close to the 97.1% on validation at 23.4%.

Banking77 cascade quality against escalation rate

Invoices (8-field JSON extraction)

Test split, n = 600. The test set is 6 layouts never seen in training, so the effective sample is 6 layouts; the template-cluster bootstrap below says how wide that makes the uncertainty. Cells are the score and its 95% paired bootstrap interval (document-level). A row aggregating more than one seed shows mean ± sample standard deviation only (see the per-seed table below for each seed's own score, and interval where the eval recorded one).

System n JSON validity Field micro-F1 Field EM Doc EM Agreement ECE AUROC
teacher (deepseek/deepseek-v4-flash-0731) 600 100.0% [100.0, 100.0] 99.5% [99.3, 99.7] 99.5% [99.3, 99.7] 96.0% [94.5, 97.5] reference — —
base 0.5B zero-shot (schema in the prompt) 600 42.3% [38.7, 46.2] 37.0% [34.3, 40.0] 24.8% [22.5, 27.4] 1.8% [0.8, 3.0] 36.9% [34.1, 39.9] 15.9% [14.2, 17.8] 0.770 [0.717, 0.824]
base 1.5B zero-shot 600 62.3% [58.2, 66.2] 73.3% [70.2, 76.1] 58.9% [54.9, 62.6] 40.3% [36.3, 44.3] 73.3% [70.2, 76.1] 9.7% [8.2, 11.5] 0.981 [0.970, 0.989]
student qwen2.5-0.5b (teacher labels, 3 seeds) 600 97.1% ± 4.3 87.2% ± 4.4 85.9% ± 4.7 39.2% ± 21.2 86.8% ± 4.5 15.6% ± 13.3 0.874 ± 0.113
student qwen2.5-1.5b (teacher labels) 600 99.2% [98.3, 99.8] 94.1% [93.3, 94.9] 94.0% [93.0, 94.9] 67.5% [63.8, 71.3] 93.5% [92.6, 94.4] 14.1% [11.9, 17.6] 0.846 [0.813, 0.879]
student qwen2.5-1.5b (gold labels) 600 95.0% [93.0, 96.7] 93.8% [92.5, 94.9] 91.4% [89.3, 93.1] 78.2% [74.8, 81.5] 93.5% [92.1, 94.6] 7.8% [5.8, 10.6] 0.902 [0.869, 0.931]
cascade (qwen2.5-1.5b-full-s13, t = 0.1524) 600 100.0% [100.0, 100.0] 94.6% [93.9, 95.3] 94.9% [94.3, 95.6] 69.2% [65.3, 72.8] 94.1% [93.3, 94.9] — —

Per-seed scores for student qwen2.5-0.5b (teacher labels, 3 seeds) (selected run: qwen2.5-0.5b-full-s13):

Run Seed JSON validity Field micro-F1 Field EM Doc EM Agreement ECE AUROC
qwen2.5-0.5b-full-s13 (validation-selected) 13 100.0% [100.0, 100.0] 90.9% [89.9, 91.8] 91.3% [90.4, 92.2] 55.3% [51.3, 59.3] 90.4% [89.3, 91.4] 8.4% [6.7, 11.8] 0.913 [0.888, 0.935]
qwen2.5-0.5b-full-s14 14 92.2% 88.5% 84.2% 47.2% 88.3% 7.3% 0.962
qwen2.5-0.5b-full-s15 15 99.2% 82.3% 82.3% 15.2% 81.8% 31.0% 0.747

Template-cluster bootstrap over 6 groups (one per layout): 95% intervals.

System JSON validity Field micro-F1 Field EM Doc EM Agreement ECE AUROC
teacher (deepseek/deepseek-v4-flash-0731) [100.0, 100.0] [98.4, 100.0] [98.5, 100.0] [88.0, 100.0] [100.0, 100.0] — —
base 0.5B zero-shot (schema in the prompt) [29.0, 55.7] [23.9, 50.3] [15.4, 36.0] [0.0, 5.2] [23.5, 50.3] [11.1, 21.3] [0.662, 0.916]
base 1.5B zero-shot [32.6, 86.8] [46.7, 89.1] [30.4, 82.3] [19.7, 58.8] [46.7, 89.1] [5.0, 14.5] [0.967, 0.993]
student qwen2.5-0.5b (teacher labels, 3 seeds) [100.0, 100.0] [82.3, 96.3] [83.2, 96.4] [30.7, 73.3] [80.8, 96.3] [5.0, 19.5] [0.873, 0.943]
student qwen2.5-1.5b (teacher labels) [97.8, 100.0] [87.1, 98.5] [86.9, 98.4] [38.7, 89.3] [85.5, 98.5] [4.2, 41.6] [0.696, 0.977]
student qwen2.5-1.5b (gold labels) [87.3, 99.0] [83.2, 98.8] [77.8, 98.3] [47.0, 94.7] [82.1, 98.8] [4.1, 28.7] [0.871, 0.991]
cascade (qwen2.5-1.5b-full-s13, t = 0.1524) [100.0, 100.0] [88.3, 98.7] [89.1, 98.7] [41.0, 89.8] [86.7, 98.7] — —
Per-template scores (student, teacher, cascade)

student:

Template n JSON validity Field micro-F1 Field EM Doc EM Agreement
email-13 100 100.0% 99.2% 99.2% 94.0% 99.2%
email-14 100 100.0% 95.9% 96.1% 69.0% 95.9%
email-15 100 100.0% 98.8% 98.9% 91.0% 98.8%
layout-13 100 99.0% 98.8% 98.4% 94.0% 98.8%
layout-14 100 96.0% 77.1% 76.9% 0.0% 73.8%
layout-15 100 100.0% 94.2% 94.5% 57.0% 94.2%

teacher:

Template n Field micro-F1 Field EM Doc EM
email-13 100 100.0% 100.0% 100.0%
email-14 100 100.0% 100.0% 100.0%
email-15 100 100.0% 100.0% 100.0%
layout-13 100 100.0% 100.0% 100.0%
layout-14 100 96.8% 97.0% 76.0%
layout-15 100 100.0% 100.0% 100.0%

cascade:

Template n JSON validity Field micro-F1 Field EM Doc EM Agreement
email-13 100 100.0% 99.2% 99.2% 94.0% 99.2%
email-14 100 100.0% 96.0% 96.2% 70.0% 96.0%
email-15 100 100.0% 98.8% 98.9% 91.0% 98.8%
layout-13 100 100.0% 99.3% 99.4% 95.0% 99.3%
layout-14 100 100.0% 79.5% 80.9% 4.0% 76.2%
layout-15 100 100.0% 94.7% 95.0% 61.0% 94.7%
Per-field exact match

Against the gold:

Field student teacher cascade
vendor_name 88.7% 96.0% 90.2%
invoice_number 83.0% 100.0% 83.8%
invoice_date 94.8% 100.0% 95.7%
due_date 88.8% 100.0% 89.8%
currency 99.2% 100.0% 100.0%
total_amount 99.2% 100.0% 100.0%
tax_amount 99.2% 100.0% 100.0%
po_number 99.2% 100.0% 100.0%

Against the teacher:

Field student cascade
vendor_name 84.7% 86.2%
invoice_number 83.0% 83.8%
invoice_date 94.8% 95.7%
due_date 88.8% 89.8%
currency 99.2% 100.0%
total_amount 99.2% 100.0%
tax_amount 99.2% 100.0%
po_number 99.2% 100.0%
Per-trait breakdown

student:

Trait n JSON validity Field micro-F1 Field EM Doc EM Agreement
(none) 91 98.9% 94.3% 93.8% 70.3% 93.4%
distractor_amounts 306 99.0% 93.9% 93.8% 67.6% 93.4%
eu_number_format 156 98.7% 95.0% 94.7% 73.1% 94.7%
label_typos 59 98.3% 94.5% 93.6% 71.2% 94.3%
missing_optional 177 99.4% 92.5% 93.6% 63.3% 91.8%
net_terms_due 93 97.8% 87.8% 87.5% 35.5% 87.1%
quoted_reply_chain 126 100.0% 94.4% 94.7% 68.3% 94.2%

teacher:

Trait n Field micro-F1 Field EM Doc EM
(none) 91 99.0% 99.0% 92.3%
distractor_amounts 306 99.6% 99.6% 96.7%
eu_number_format 156 99.7% 99.7% 97.4%
label_typos 59 99.8% 99.8% 98.3%
missing_optional 177 99.4% 99.5% 96.0%
net_terms_due 93 99.3% 99.3% 94.6%
quoted_reply_chain 126 99.8% 99.8% 98.4%

cascade:

Trait n JSON validity Field micro-F1 Field EM Doc EM Agreement
(none) 91 100.0% 94.9% 94.9% 71.4% 94.0%
distractor_amounts 306 100.0% 94.5% 94.8% 69.0% 94.0%
eu_number_format 156 100.0% 95.8% 96.1% 75.0% 95.5%
label_typos 59 100.0% 95.1% 95.3% 72.9% 94.8%
missing_optional 177 100.0% 92.9% 94.3% 65.0% 92.3%
net_terms_due 93 100.0% 89.6% 90.2% 41.9% 88.9%
quoted_reply_chain 126 100.0% 94.4% 94.7% 68.3% 94.2%
Value
Target Agreement ≥ 97.0% against the teacher
Chosen threshold (on validation) 0.1524
Escalation rate, valid / test 0.0% / 1.7%
Cascade − teacher, Agreement against the teacher (target) -5.9 pts [-6.7, -5.1], cluster [-13.3, -1.3]
Cascade − teacher, Field micro-F1 (gold) -4.8 pts [-5.5, -4.2], cluster [-10.1, -1.3]
Target held on test no

Target held on test: no — Agreement against the teacher met the 97.0% target on validation (98.5% at 0.0% escalation) but fell to 94.1% on test (1.7% escalation).

Invoices cascade quality against escalation rate

Cost and latency

Energy assumptions: local cost = 20 W x wall time x $0.3/kWh; hardware amortisation off.

Banking77

Measured on a Mac Studio, Apple M4 Max, 128 GB (Mac16,9), macOS 26.5.2, 2026-09-27:

System $/1k recorded $/1k list price p50 ms p95 ms Source
Teacher only $0.00593 $0.0581 594 1,242 recorded live labelling calls at concurrency 8 (cache hits and retries excluded)
Student only $0.0000351 — 19.9 31.7 bench against serve --threshold 0
Cascade $0.00136 $0.0129 21.5 854 composed per request, 22.0% escalated

For comparison, the MacBook Air's student latency was 44.5 ms p50 / 68.2 ms p95 (in-process eval, not through the server).

Break-even at 17,419 requests at the recorded teacher cost, 1,762 requests at the list price without prompt caching.

Invoices

Measured on a Mac Studio, Apple M4 Max, 128 GB (Mac16,9), macOS 26.5.2, 2026-09-27:

System $/1k recorded $/1k list price p50 ms p95 ms Source
Teacher only $0.0462 $0.0577 1,354 2,427 recorded live labelling calls at concurrency 8 (cache hits and retries excluded)
Student only $0.000729 — 441 498 bench against serve --threshold 0
Cascade $0.00119 $0.00130 442 499 composed per request, 1.0% escalated

For comparison, the MacBook Air's student latency was 1,391 ms p50 / 1,988 ms p95 (in-process eval, not through the server).

Break-even at 3,140 requests at the recorded teacher cost, 2,505 requests at the list price without prompt caching.

The Why section above states the teacher answered in about 594 ms at p50 and 1,242 ms at p95 (Banking77 labelling calls, 2026-09-26); that is this same recorded Banking77 teacher latency.

Live bench cross-check

Banking77 (measured on a Mac Studio, Apple M4 Max, 128 GB (Mac16,9), macOS 26.5.2, 2026-09-27)

Mode Run Date n p50 ms p95 ms Escalated Spend Load average
Student only qwen2.5-0.5b-full-s13 2026-09-27T02:51:24+00:00 300 19.9 31.7 0.0% $0 15.01/12.64/10.21
Cascade qwen2.5-0.5b-full-s13 2026-09-27T02:52:20+00:00 300 31.3 815 21.3% $0.000500 14.16/12.56/10.22

Composed (from the test split) vs measured: p50 21.5 vs 31.3 ms, p95 854 vs 815 ms.

Invoices (measured on a Mac Studio, Apple M4 Max, 128 GB (Mac16,9), macOS 26.5.2, 2026-09-27)

Mode Run Date n p50 ms p95 ms Escalated Spend Load average
Student only qwen2.5-1.5b-full-s13 2026-09-27T02:54:45+00:00 300 441 498 0.0% $0 11.04/11.93/10.15
Cascade qwen2.5-1.5b-full-s13 2026-09-27T02:57:14+00:00 300 445 505 1.3% $0.000199 9.45/10.91/10.00

Composed (from the test split) vs measured: p50 442 vs 445 ms, p95 499 vs 505 ms.

Training on the MacBook Air

Banking77

Run Base Examples Dropped (length) Iterations Epochs Wall min Peak GB Tokens/s Adapter MB
qwen2.5-0.5b-full-s13 mlx-community/Qwen2.5-0.5B-Instruct-4bit 8,874 0 2,219 2.00 25.1 3.45 668 35.2
qwen2.5-1.5b-full-s13 mlx-community/Qwen2.5-1.5B-Instruct-4bit 8,874 0 2,219 2.00 59.1 6.16 280 73.9

Invoices

Run Base Examples Dropped (length) Iterations Epochs Wall min Peak GB Tokens/s Adapter MB
qwen2.5-0.5b-full-s13 mlx-community/Qwen2.5-0.5B-Instruct-4bit 2,000 0 500 2.00 34.1 4.17 792 35.2
qwen2.5-1.5b-full-s13 mlx-community/Qwen2.5-1.5B-Instruct-4bit 2,000 0 500 2.00 96.7 5.20 276 73.9

Load average recorded alongside these runs ranged up to 4.27. The MacBook Air is fanless and can throttle under sustained load; that is why the load average is recorded next to every timing rather than assumed away.

Reproducibility on a second machine

scripts/reproduce.sh (full profile) was run again on a second machine and compared with scripts/compare_reports.py; tolerance 2.0 points on each test-split metric.

Reference: MacBook Air, Apple M5, 24 GB (Mac17,4), macOS 26.6.2, 2026-09-27. Rerun: Mac Studio, Apple M4 Max, 128 GB (Mac16,9), macOS 26.5.2, 2026-09-27.

Task Max abs difference All rows within tolerance
Banking77 0.78 pts yes
Invoices 10.17 pts no
Task Row Metric Reference Rerun Diff
Banking77 teacher (deepseek/deepseek-v4.1-flash) Accuracy 75.8% 75.8% +0.00 pts
Banking77 TF-IDF + logistic regression (teacher labels) Accuracy 72.5% 72.5% +0.00 pts
Banking77 student qwen2.5-0.5b (teacher labels, 3 seeds) Accuracy 75.7% 75.9% +0.25 pts
Banking77 student qwen2.5-1.5b (teacher labels) Accuracy 76.5% 76.3% -0.16 pts
Invoices teacher (deepseek/deepseek-v4-flash-0731) Field micro-F1 99.5% 99.5% +0.00 pts
Invoices student qwen2.5-0.5b (teacher labels, 3 seeds) Field micro-F1 87.2% 88.2% +0.96 pts
Invoices student qwen2.5-1.5b (teacher labels) Field micro-F1 94.1% 94.5% +0.46 pts
Invoices student qwen2.5-1.5b (gold labels) Field micro-F1 93.8% 91.5% -2.37 pts

Banking77: the selected run differed — qwen2.5-0.5b-full-s13 on the reference machine vs qwen2.5-1.5b-full-s13 on the rerun (reference: “large gains 0.97 points, below the 1.00-point minimum”; rerun: “large gains 1.26 points (>= 1.00) at 1.87x the small p95 (< 3x)”). Invoices: the selected run matched on both machines (qwen2.5-1.5b-full-s13).

All rows within tolerance across every task: no.

Outside the tolerance: Invoices student qwen2.5-0.5b (teacher labels, 3 seeds), Doc EM: 39.2% vs 41.6% (+2.33 points); Invoices student qwen2.5-1.5b (teacher labels), Doc EM: 67.5% vs 74.3% (+6.83 points); Invoices student qwen2.5-1.5b (gold labels), JSON validity: 95.0% vs 98.2% (+3.17 points); Invoices student qwen2.5-1.5b (gold labels), Field micro-F1: 93.8% vs 91.5% (-2.37 points); Invoices student qwen2.5-1.5b (gold labels), Doc EM: 78.2% vs 68.0% (-10.17 points); Invoices student qwen2.5-1.5b (gold labels), Agreement: 93.5% vs 91.0% (-2.43 points).

Calibration

Banking77 reliability diagram, raw vs isotonic

Two confidence definitions were compared on validation, before either was used at test time: the primary is the trie-constrained greedy label's own renormalised token-probability product; the alternative is free greedy generation scored by the mean per-token log-probability. AUROC is against the teacher and the gold label; ECE against each.

Task Split Confidence n AUROC/teacher AUROC/gold ECE/teacher ECE/gold
Banking77 valid primary (chosen) 1,030 0.877 0.799 2.8% 17.0%
Banking77 valid alternative 1,030 0.874 0.811 9.2% 24.1%
Banking77 test primary (chosen) 3,075 0.903 0.799 3.0% 15.0%
Banking77 test alternative 3,075 0.901 0.807 9.3% 21.8%
Invoices valid primary (chosen) 400 0.904 0.913 3.7% 4.3%
Invoices valid alternative 400 0.903 0.908 10.0% 9.3%
Invoices test primary (chosen) 600 0.846 0.846 14.1% 14.1%
Invoices test alternative 600 0.858 0.858 32.1% 32.1%

Banking77 student qwen2.5-0.5b-full-s13, ECE on test vs gold: 15.0% raw vs 3.7% after isotonic calibration.

Invoices student qwen2.5-1.5b-full-s13, ECE on test vs gold: 14.1% raw vs 15.6% after isotonic calibration.

Choosing the base model

Banking77

Base model Validation metric p95 ms
mlx-community/Qwen2.5-0.5B-Instruct-4bit 88.2% 74.8
mlx-community/Qwen2.5-1.5B-Instruct-4bit 89.1% 163

Rule: use the larger base only if it gains at least 1 point on the validation metric and its p95 latency stays under 3.0x the smaller model's. Here the gain is below the 1 point minimum (+0.97 points) and the p95 ratio (2.18x) is under the 3.0x limit, so the smaller base was kept.

Invoices

Base model Validation metric p95 ms
mlx-community/Qwen2.5-0.5B-Instruct-4bit 94.7% 1,102
mlx-community/Qwen2.5-1.5B-Instruct-4bit 98.5% 1,974

Rule: use the larger base only if it gains at least 1 point on the validation metric and its p95 latency stays under 3.0x the smaller model's. Here the gain meets the 1 point minimum (+3.79 points) and the p95 ratio (1.79x) is under the 3.0x limit, so the larger base was selected.

The rule compares point estimates on the validation split (the best seed of each base); no interval is computed for this decision, so a gain close to the minimum can go either way on a rerun (see Reproducibility above).

What didn't work

  • Learning-rate schedule at the spec's peak rate (quick profile, 200 iterations, 3 seeds, Banking77): on the MacBook Air (Apple M5) a constant rate reached 32.1% ± 31.8% mean validation agreement with the teacher and linear warm-up + cosine decay 68.9% ± 8.3%, with 1 and 0 of 3 seeds diverging (agreement below 10%); on the Mac Studio (Apple M4 Max) a constant rate reached 23.3% ± 38.0% mean validation agreement with the teacher and linear warm-up + cosine decay 33.9% ± 32.9%, with 2 and 1 of 3 seeds diverging (agreement below 10%). Warm-up + cosine is the default and did better on average, but it does not make short runs at this peak rate reliable. In the full profile, Banking77 converged on every seed; on invoices, the best validation checkpoint of 1 of 3 on the MacBook Air (Apple M5) and 1 of 3 on the Mac Studio (Apple M4 Max) 0.5B seeds came at or before the end of warm-up, so those students stopped early.

  • The bigger Banking77 student (1.5B vs 0.5B, teacher labels) gained +0.97 points on validation agreement, below the 1 point minimum the base-model rule requires, for 2.18x the p95 latency (limit 3.0x); the rule kept the 0.5B base.

  • A Banking77 student trained on the gold labels reached 92.0% accuracy against gold, above both the same student trained on teacher labels (75.5%) and the teacher itself (75.8%); part of that gap is Banking77's own label noise (Ying and Thomas, 2022, flag about 14% of the training utterances as potential label errors), which caps how high any model's accuracy against gold can go.

  • Teacher prompt variants (same 200 Banking77 validation queries, same teacher model): snake-case labels (+0.5 pts accuracy at 1.3x the prompt tokens, 3.3x the cost per 1k); labels with examples (+4.0 pts accuracy at 4.3x the prompt tokens, 2.9x the cost per 1k). The spec prompt was kept for the recorded run.

  • Two confidence definitions were compared on validation, before either was used at test time: Banking77 AUROC 0.877 vs 0.874 (about equal), ECE 2.8% vs 9.2%, about 3.3x lower for the primary; Invoices AUROC 0.904 vs 0.903 (about equal), ECE 3.7% vs 10.0%, about 2.7x lower for the primary; so the primary (trie-constrained token-probability product) was kept over the alternative (free greedy generation, mean per-token log-probability). Reported honestly, on test: Banking77 AUROC 0.903 vs 0.901 (about equal), ECE 3.0% vs 9.3% (lower for the primary); Invoices AUROC 0.846 vs 0.858 (higher for the alternative), ECE 14.1% vs 32.1% (lower for the primary).

  • Invoices: the threshold chosen on the validation layouts did not transfer to the unseen test layouts — Agreement against the teacher met the 97.0% target on validation (98.5%, 0.0% escalation) but fell to 94.1% on test (1.7% escalation).

Spend and downloads

Task Teacher labelling Bake-off Live bench
Banking77 $0.0796 $0.0211 $0.000500
Invoices $0.138 $0.0116 $0.000199

Total spend for the whole build: $0.251 of a $14.00 global cap (17,302 teacher calls).

Model downloads: 1.17 GB, within the 1.5 GB budget (student base 0.5B: 0.29 GB; student base 1.5B: 0.88 GB).

Configuration reference

A task lives in tasks/<task>/task.yaml next to its teacher prompt and its labels file or JSON Schema. String values can use ${NAME} or ${NAME:-default} environment references. Unknown keys are errors, and every validation error names the offending key.

Key Default Meaning
task (required) Task name: 1–64 letters, digits, ., _ or -. Data lives in $TASKDISTILL_HOME/<task>/.
type (required) classification or extraction.
labels_file — Classification: one canonical label per line (required).
schema_file — Extraction: a JSON Schema with type: object and properties (required).
input.from last_user_message Where the task input is in each request: the last user message, or regex.
input.regex null With from: regex: a pattern over the last user message with a named group input. Requests that do not match go to the teacher (input_unparsed).
teacher.base_url https://openrouter.ai/api/v1 Any OpenAI-compatible endpoint.
teacher.api_key_env TASKDISTILL_TEACHER_API_KEY Environment variable holding the teacher key; falls back to OPENROUTER_API_KEY.
teacher.model (required) The slug your application calls today.
teacher.prompt_file teacher_prompt.md The teacher's system prompt, sent verbatim.
teacher.temperature 0 Teacher sampling temperature.
teacher.max_tokens 256 Teacher completion cap; outputs that hit it are flagged as truncated.
teacher.response_format null For example {type: json_object}, when the provider supports it.
teacher.extra_body {} Extra request fields: provider pinning (provider: {order: [...], allow_fallbacks: false}), reasoning controls. Recorded in the replay manifest, not in the request key.
student.base_model mlx-community/Qwen2.5-0.5B-Instruct-4bit Hugging Face id or local directory of the student base; --base overrides it.
student.system_prompt (required) The student's short system prompt, used in training and serving.
student.max_tokens 16 Student generation cap.
student.response_template null Renders student answers and canonical escalations, e.g. '{"intent": "{label}"}' (literal replacement of {label}).
train.profile full quick caps training at 200 iterations and validates less often; --profile overrides it.
train.lora_rank 16 LoRA rank (scale 20, as mlx-lm's default).
train.lora_layers all Number of top transformer blocks to adapt, or all.
train.learning_rate 1.0e-4 Peak learning rate (Adam).
train.batch_size 8 Examples per step.
train.epochs 2 Passes over the training split.
train.max_seq_len 512 Longer examples are dropped by curate and training, never truncated.
train.seed 13 Training seed (--seed overrides it).
train.lr_schedule warmup_cosine Linear warm-up over 10% of the iterations (at most 100), then cosine decay to 10% of the peak; or constant.
curate.dedupe.exact true Exact duplicates on the normalised input (NFKC, whitespace-collapsed, case-folded).
curate.dedupe.near_dup_jaccard 0.9 MinHash LSH threshold on word 3-shingles (0.5–1). Inputs under 3 words are compared exactly.
curate.pii.enabled true Regex scrub with checksums before dedupe; matches become <EMAIL>, <PHONE>, <IBAN>, <CARD>, <IPV4>, <SSN>.
curate.pii.kinds all six Any subset of email, phone, iban, card, ipv4, ssn.
curate.split.predefined meta.split Meta field whose value (train, valid, test) fixes a row's split; null to ignore.
curate.split.val, curate.split.test 0.1, 0.1 Fractions for rows without a predefined split.
curate.split.stratify true Stratify the random split by label (classification).
curate.split.group_by null Meta field whose groups stay in one split (for example meta.template); also the clusters of the cluster bootstrap. Without it, eval uses meta.group when present.
curate.split.seed 13 Split seed.
cascade.reference teacher What the target is measured against: teacher (works with no gold labels) or gold.
cascade.metric agreement agreement (label equality, or field micro-F1 against the teacher for extraction), accuracy or macro_f1 (classification), field_f1 (extraction).
cascade.target — Minimum cascade quality on validation, e.g. 0.97. Set exactly one of target and max_drop.
cascade.max_drop — Maximum drop below the teacher's own score on the reference, usually with reference: gold.
cascade.on_teacher_error student On a teacher timeout or error: return the student's answer (student-fallback) or an HTTP 502 (error). A replay miss is always an error. serve's escalation gives up quickly rather than retrying like a batch job: a 10 s HTTP timeout, at most 1 retry, Retry-After honoured up to 2 s, all within a 30 s deadline for the whole call (wait for a free connection included), so a hung or rate-limited teacher falls back within seconds, not minutes.
cascade.escalation_response canonical canonical normalises and renders the teacher's answer like a student answer; raw returns it verbatim.
cost.local_watts 20 Power draw assumed for local inference and training.
cost.usd_per_kwh 0.30 Electricity price.
cost.hardware_usd, cost.amortisation_hours 0, 0 Hardware amortisation, applied only when both are above 0.
budget.usd_cap null Optional per-task spend cap; TASKDISTILL_BUDGET_USD always applies.

Environment (see .env.example): TASKDISTILL_TEACHER_BASE_URL, TASKDISTILL_TEACHER_API_KEY (falls back to OPENROUTER_API_KEY), TASKDISTILL_TEACHER_MODEL, TASKDISTILL_BUDGET_USD (global spend cap, default 5.00), TASKDISTILL_SERVER_TOKEN (required to bind capture or serve beyond localhost) and TASKDISTILL_HOME (the workspace: store, response cache, ledger, curated data and runs; default ./.taskdistill).

Spend control. Every command that can call the teacher takes --max-usd (a cap for that run) and --yes. Before each call the ledger reserves the worst case (UTF-8 bytes of the messages plus 16 per message as prompt tokens, plus max_tokens of completion) and settles it with the usage.cost the API returns, or with the token counts times the pricing snapshot when the response has no cost, so concurrent calls can never cross a cap. Batches first run a 50-request sample and need --yes when the projection exceeds $0.50. taskdistill budget prints the spend by task and phase.

Limitations

  • Platforms. The MLX path needs Apple Silicon and macOS 14 or later (the minimum of the pinned mlx wheels). On other machines train --backend mlx stops with a message pointing to --backend torch. The torch path (transformers + PEFT, Unsloth when importable with CUDA) is tested on CPU with a tiny model in CI; it has not been run on a GPU, and the Unsloth branch is covered only by a dispatch test with a mock.
  • PII scrub is best-effort. It is a regex scrub with checksums for e-mail addresses, phone numbers, IBANs, card numbers, IPv4 addresses and US SSNs. It does not find names, street addresses or anything else, and it is not an anonymisation guarantee. It also creates a train/serve skew: the student trains on placeholders such as <EMAIL> but sees raw text when it serves.
  • Training on Metal is not bit-for-bit deterministic. Two runs with the same seed can differ slightly; the results report the spread over seeds and the tolerance used when checking that README numbers reproduce.
  • Scope. Single-turn classification and JSON extraction only: no summaries or chat replies, no multi-turn or tool-calling tasks, no multi-label classification. Streamed responses are forwarded but not captured. The server generates one request at a time (no batching) and serves one task per process. Escalations for unsupported or unparsable requests are returned verbatim; only low-confidence escalations are normalised.
  • Evaluation limits. The invoice test set is 600 documents from 6 layouts never seen in training, so its effective sample is 6 layouts; the template-cluster bootstrap intervals say how wide that makes the uncertainty. The invoices are synthetic and cleaner than real mail. Banking77's gold labels are noisy (Ying and Thomas, 2022, flag about 14% of the training utterances as potential label errors), which caps any model's measured accuracy against gold.
  • Cost figures are estimates where they are not measured. Local cost is an energy estimate (20 W at $0.30/kWh by default, both configurable), not a power measurement. With an inexpensive teacher whose provider caches the shared prompt, the money saved per request is small; the main gains are latency and keeping inputs on the machine.
  • Latency numbers are dated measurements from two machines, and each table names its machine. Training and the in-process student latency ran on a fanless MacBook Air (Apple M5, 24 GB) that throttles under sustained load and is shared with other work; the cost and latency tables, the live bench and the break-even volumes come from a Mac Studio (Apple M4 Max, 128 GB) used only for those runs. The load average is recorded next to every timing, and live numbers are not expected to reproduce exactly on other hardware or under other load.
  • Adapters are merged into the 4-bit base in memory for evaluation and serving, which avoids the extra LoRA matrix multiplications an unmerged adapter runs at every decoding step. The re-quantisation shifts confidences slightly, which is why the threshold is chosen on the same merged model that serves.
  • Teacher terms. Some commercial APIs forbid using their outputs to train other models. Check your provider's terms before distilling its outputs; see ADR 5 for the checks made here.

Roadmap

  • Shadow mode and canary routing for the cascade, so a new student can be compared on live traffic before it answers.
  • Scheduled retraining and active learning from escalated requests (the served log already records them).
  • Drift alerts on the observed escalation rate (today report --from-serve-log shows drift; nothing alerts).
  • Capture of streamed responses, multi-turn and tool-calling tasks.
  • Constrained JSON decoding for extraction, and request batching in the server.

Data and licences

  • taskdistill is Apache-2.0 (Copyright 2026 Andrii Boiko). The synthetic invoice generator and its output are part of the repository and share that licence. A 20-document sample of the generated invoices (one per training layout, covering every trait, with their gold fields) is in examples/invoices/; scripts/make_examples.py regenerates it.

  • Banking77 (Casanueva et al., 2020) is licensed under CC BY 4.0. The demo downloads the CSVs at runtime from a pinned commit of PolyAI-LDN/task-specific-datasets and checks their SHA-256. If GitHub is unavailable it falls back to the Parquet mirror legacy-datasets/banking77 on the Hugging Face Hub (install the hub extra); PolyAI/banking77 itself is a script-based Hub repository with no Parquet conversion. This repository does not redistribute the CSVs: it ships only the teacher's outputs keyed by request hash. Modifications: the official training set is split 90/10 into train and validation (stratified, seed 13), and duplicates are removed as described above.

    @inproceedings{casanueva-etal-2020-efficient,
        title = "Efficient Intent Detection with Dual Sentence Encoders",
        author = "Casanueva, I{\~n}igo and Tem{\v{c}}inas, Tadas and Gerz, Daniela and Henderson, Matthew and Vuli{\'c}, Ivan",
        booktitle = "Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI",
        year = "2020",
        publisher = "Association for Computational Linguistics",
        url = "https://aclanthology.org/2020.nlp4convai-1.5/",
        doi = "10.18653/v1/2020.nlp4convai-1.5",
        pages = "38--45"
    }
    
  • Student bases: Qwen2.5-0.5B-Instruct and Qwen2.5-1.5B-Instruct are Apache-2.0; the demos use the 4-bit MLX conversions mlx-community/Qwen2.5-0.5B-Instruct-4bit and mlx-community/Qwen2.5-1.5B-Instruct-4bit at pinned revisions. No trained adapters are published.

  • Teacher outputs: DeepSeek V4.1 Flash (Banking77) and DeepSeek V4 Flash 0731 (invoices) through OpenRouter, pinned to the DeepInfra provider. The DeepSeek V4 weights are MIT-licensed, DeepSeek's platform terms allow training other models on outputs, and DeepInfra's terms do not restrict the use of outputs (ADR 5). The outputs were recorded on 2026-09-26.

Licence

Apache-2.0. See LICENSE. Copyright 2026 Andrii Boiko.

Release files for taskdistill 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for taskdistill 0.1.1
File Size Uploaded
taskdistill-0.1.1.tar.gz 2.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for taskdistill 0.1.1
File Interpreter ABI Platform
taskdistill-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 3.8 MB

Release files / taskdistill-0.1.1.tar.gz

Download URL taskdistill-0.1.1.tar.gz
Size 2.1 MB
Tags Source
SHA-256 checksum
How to use checksums
3c77950d5983ebfdabc3d68005a3eb2413be66958bdc23380fa30d32ae489e4e
BLAKE2b-256 checksum
How to use checksums
91ea84e417fddee65553372d446249d13a0d4d59f3db7c0a13b2d69587ac667b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / taskdistill-0.1.1-py3-none-any.whl

Download URL taskdistill-0.1.1-py3-none-any.whl
Size 1.7 MB
Tags Python 3
SHA-256 checksum
How to use checksums
e279bd5999e295d22bc475822b325b4457de25a6a4075b63e862767ed0061eac
BLAKE2b-256 checksum
How to use checksums
5bee163400c40618781f2ba92f3bd141b9412c46b82610b607bb0ed521a5c6c6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page