Skip to main content

Loggetta

Plan the run. Train the model. Keep the evidence.

Loggetta checks your hardware, chooses a supported configuration for a Mixture-of-Experts (MoE) model, and runs QLoRA fine-tuning. It saves the plan and a JSON run report with memory use, timing, and checks that the selected optimizations actually ran. One install includes the runtime (experts4bit-qlora) and the GPU kernels (grouped-nf4-gemm).

pip install loggetta
loggetta inspect
loggetta plan Qwen/Qwen3-30B-A3B --seq 2048

The plan shows where weights will live, estimated memory use, and why alternatives were rejected. It checks the budget before downloading model weights. Estimates can miss; they are not an out-of-memory guarantee.

New in 0.4.0

  • Dense models: planned, not supported for training. Dense plans are estimates checked in sample; out of sample (DQ7) they missed, and the DQ10 reading is pending. Dense execution needs --allow-development-executor until DQ8's 24 GB reading passes. See Dense training.
  • The DQ10 reserve policy is opt-in, for registered dense runs. Defaults do not change.
  • MoE training estimates price the grouped_nf4 backward pass. OLMoE plans on an RTX A2000 were about 0.2 GiB short; in sample they now cover the measured peak. Near a budget, a plan may pick the reference kernel or host residency where it picked resident grouped_nf4 before.
  • Single-stream serve plans name the speed-ups a default server runs for the model's family, and quote a decode speed only from a measured run of the same setup.
  • Fix: resident grouped_nf4 training runs again with experts4bit-qlora 0.49.0 or later. Loggetta 0.3.x stopped before the first step (#47); the workaround was E4B_ABSMAX_DQ=0.

Full list: CHANGELOG.

Train on your data

loggetta train Qwen/Qwen3-30B-A3B \
  --dataset ./data/train.jsonl --format text \
  --seq 512 --micro-batch 1 --steps 20 --seed 42 \
  --out runs/my-training --adapter-out adapters/my-adapter

Local JSONL, JSON, CSV, Parquet and TXT files, Hub datasets, Alpaca instructions and text-only chats are supported. Data is validated and tokenized before weights load. Chat data trains only the assistant turns by default. Reload the adapter with loggetta.load_adapter("adapters/my-adapter"). See the training guide.

To run a saved decision later: loggetta plan MODEL --out plan.json, then loggetta execute plan.json --out runs/. Pass earlier run reports back with --observations runs/ and later plans use those measurements.

Measured results

The included runtime and kernels do the compute; these are matched training runs, not planner benchmarks.

Workload Result
Qwen3-30B-A3B QLoRA · RTX 5090 Unsloth spends 1.92× e4b's GPU time per step, and 2.80× its wall-clock time on an AMD EPYC 7713 host. Comparable held-out loss; Unsloth peaked lower (24.27 vs 26.16 GB). Result
Why two numbers GPU time doesn't depend on the host. Unsloth runs about 14× e4b's CPU operations per step, so its wall-clock time grows on a slower host. Earlier wall-clock readings, before e4b's current defaults: 2.352× and 2.468×.
Planner memory check · Qwen3-30B-A3B · RTX 5090 24.54 GiB estimated process peak, 24.34 GiB measured, after calibration from earlier runs. One in-sample case, not a guarantee. In the MoE audit re-run after 0.4.0's grouped_nf4 term (evidence/2026-10-09-moe-plan-vs-driver-after-gnf4), one RTX A2000 training plan is still under its measured peak (1.014×), where it borrows another model's reserve. Plan vs run · MoE audit

The comparison used torch 2.12.1+cu130 and transformers 5.5.0 for both frameworks, with matched adapters, initialization and tokens. Loggetta does not predict throughput.

Models

Model Tested in Loggetta
OLMoE-1B-7B-0924, Granite-3.1-3B-A800M training and serving run
Granite-4.0-H-tiny training run; serving refused (Mamba state)
Qwen3-30B-A3B, Qwen3.6-35B-A3B, ERNIE-4.5-21B-A3B planned; serving validated
LFM2-8B-A1B planned; serving refused (conv state)
Mixtral-8x7B-Instruct planned
Gemma-4-26B-A4B-it, Nemotron-3.5-Lightning-30B-A3B supported by the runtime; not yet in Loggetta's sweep

Every row is supported by the included runtime's QLoRA path. Hybrid models (Qwen3.6, Granite-4.0-H, LFM2, Nemotron-H) need --packing concat with chat or Alpaca data. The runtime's capability register is the authority.

Scope

Single-GPU MoE planning and QLoRA training; serving placement is planned, not launched. Not yet: multi-GPU, throughput prediction. GPU runs need Linux, an NVIDIA CUDA GPU and a compatible PyTorch. Pre-1.0.

Dense models are not supported for training yet: their plans are estimates checked in sample (DQ7 missed out of sample, DQ10 pending), and their training runs only behind --allow-development-executor until DQ8's 24 GB reading.

GitHub · Results · Architecture · Research and releases

Metadata

Release files for loggetta 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for loggetta 0.4.0
File Size Uploaded
loggetta-0.4.0.tar.gz 151.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for loggetta 0.4.0
File Interpreter ABI Platform
loggetta-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 256.3 kB

Release files / loggetta-0.4.0.tar.gz

Download URL loggetta-0.4.0.tar.gz
Size 151.4 kB
Tags Source
SHA-256 checksum
How to use checksums
71db36b288ef723bc7c69164ad9197febadd115c7bdb110d97193d2009af8910
BLAKE2b-256 checksum
How to use checksums
7083e742754a0da2f6fa7c47c747e164a9c0a474462ee2312925fb776e0fe86d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / loggetta-0.4.0-py3-none-any.whl

Download URL loggetta-0.4.0-py3-none-any.whl
Size 104.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
06ace6dc0b1fa02e2e59917c16e6e021e599af1128b68f3b896eda3429c19dc1
BLAKE2b-256 checksum
How to use checksums
da4fe569b785426a5bb8a96d28ed62d197cdeec817cc032685a7df6c0a7f1fab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

0.5.0

2 release files

This release

0.4.0 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page