Loggetta
Plan the run. Train the model. Keep the evidence.
Loggetta checks your hardware, chooses a supported configuration for a Mixture-of-Experts (MoE) model, and runs QLoRA fine-tuning. It saves the plan and a JSON run report with memory use, timing, and checks that the selected optimizations actually ran. One install includes the runtime (experts4bit-qlora) and the GPU kernels (grouped-nf4-gemm).
pip install loggetta
loggetta inspect
loggetta plan Qwen/Qwen3-30B-A3B --seq 2048
The plan shows where weights will live, estimated memory use, and why alternatives were rejected. It checks the budget before downloading model weights. Estimates can miss; they are not an out-of-memory guarantee.
New in 0.4.0
- Dense models: planned, not supported for training. Dense plans are estimates checked in sample; out of sample
(DQ7) they missed, and the DQ10 reading is pending. Dense execution needs
--allow-development-executoruntil DQ8's 24 GB reading passes. See Dense training. - The DQ10 reserve policy is opt-in, for registered dense runs. Defaults do not change.
- MoE training estimates price the
grouped_nf4backward pass. OLMoE plans on an RTX A2000 were about 0.2 GiB short; in sample they now cover the measured peak. Near a budget, a plan may pick the reference kernel or host residency where it picked residentgrouped_nf4before. - Single-stream serve plans name the speed-ups a default server runs for the model's family, and quote a decode speed only from a measured run of the same setup.
- Fix: resident
grouped_nf4training runs again with experts4bit-qlora 0.49.0 or later. Loggetta 0.3.x stopped before the first step (#47); the workaround wasE4B_ABSMAX_DQ=0.
Full list: CHANGELOG.
Train on your data
loggetta train Qwen/Qwen3-30B-A3B \
--dataset ./data/train.jsonl --format text \
--seq 512 --micro-batch 1 --steps 20 --seed 42 \
--out runs/my-training --adapter-out adapters/my-adapter
Local JSONL, JSON, CSV, Parquet and TXT files, Hub datasets, Alpaca instructions and text-only chats are supported.
Data is validated and tokenized before weights load. Chat data trains only the assistant turns by default. Reload the
adapter with loggetta.load_adapter("adapters/my-adapter"). See the
training guide.
To run a saved decision later: loggetta plan MODEL --out plan.json, then loggetta execute plan.json --out runs/.
Pass earlier run reports back with --observations runs/ and later plans use those measurements.
Measured results
The included runtime and kernels do the compute; these are matched training runs, not planner benchmarks.
| Workload | Result |
|---|---|
| Qwen3-30B-A3B QLoRA · RTX 5090 | Unsloth spends 1.92× e4b's GPU time per step, and 2.80× its wall-clock time on an AMD EPYC 7713 host. Comparable held-out loss; Unsloth peaked lower (24.27 vs 26.16 GB). Result |
| Why two numbers | GPU time doesn't depend on the host. Unsloth runs about 14× e4b's CPU operations per step, so its wall-clock time grows on a slower host. Earlier wall-clock readings, before e4b's current defaults: 2.352× and 2.468×. |
| Planner memory check · Qwen3-30B-A3B · RTX 5090 | 24.54 GiB estimated process peak, 24.34 GiB measured, after calibration from earlier runs. One in-sample case, not a guarantee. In the MoE audit re-run after 0.4.0's grouped_nf4 term (evidence/2026-10-09-moe-plan-vs-driver-after-gnf4), one RTX A2000 training plan is still under its measured peak (1.014×), where it borrows another model's reserve. Plan vs run · MoE audit |
The comparison used torch 2.12.1+cu130 and transformers 5.5.0 for both frameworks, with matched adapters, initialization and tokens. Loggetta does not predict throughput.
Models
| Model | Tested in Loggetta |
|---|---|
| OLMoE-1B-7B-0924, Granite-3.1-3B-A800M | training and serving run |
| Granite-4.0-H-tiny | training run; serving refused (Mamba state) |
| Qwen3-30B-A3B, Qwen3.6-35B-A3B, ERNIE-4.5-21B-A3B | planned; serving validated |
| LFM2-8B-A1B | planned; serving refused (conv state) |
| Mixtral-8x7B-Instruct | planned |
| Gemma-4-26B-A4B-it, Nemotron-3.5-Lightning-30B-A3B | supported by the runtime; not yet in Loggetta's sweep |
Every row is supported by the included runtime's QLoRA path. Hybrid models (Qwen3.6, Granite-4.0-H, LFM2, Nemotron-H)
need --packing concat with chat or Alpaca data. The runtime's
capability register is the
authority.
Scope
Single-GPU MoE planning and QLoRA training; serving placement is planned, not launched. Not yet: multi-GPU, throughput prediction. GPU runs need Linux, an NVIDIA CUDA GPU and a compatible PyTorch. Pre-1.0.
Dense models are not supported for training yet: their plans are estimates checked in sample (DQ7 missed out of
sample, DQ10 pending), and their training runs only behind --allow-development-executor until DQ8's 24 GB reading.
Metadata
Release files for loggetta 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| loggetta-0.4.0.tar.gz | 151.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| loggetta-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 256.3 kB
Release files / loggetta-0.4.0.tar.gz
| Download URL | loggetta-0.4.0.tar.gz |
|---|---|
| Size | 151.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
71db36b288ef723bc7c69164ad9197febadd115c7bdb110d97193d2009af8910
|
|
BLAKE2b-256 checksum How to use checksums |
7083e742754a0da2f6fa7c47c747e164a9c0a474462ee2312925fb776e0fe86d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency logRelease files / loggetta-0.4.0-py3-none-any.whl
| Download URL | loggetta-0.4.0-py3-none-any.whl |
|---|---|
| Size | 104.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
06ace6dc0b1fa02e2e59917c16e6e021e599af1128b68f3b896eda3429c19dc1
|
|
BLAKE2b-256 checksum How to use checksums |
da4fe569b785426a5bb8a96d28ed62d197cdeec817cc032685a7df6c0a7f1fab
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency log