Skip to main content

Loggetta

One install for planning and running single-GPU MoE QLoRA workloads.

pip install loggetta

That installs Loggetta plus its current runtime and kernel stack:

  • Loggetta — hardware inventory, planning, ExecutionPlan, execution orchestration, and ExecutionReceipt
  • experts4bit-qlora — model loading, QLoRA, training, serving mechanisms, and CPU/NVMe offload
  • grouped-nf4-gemm — packed low-bit kernels and residency primitives

Loggetta examines the model, machine, workload, and constraints before loading model weights. It chooses a supported setup, explains why alternatives lost, and refuses configurations estimated not to fit. For supported training plans it dispatches into the included runtime and records what actually happened.

loggetta inspect
loggetta plan Qwen/Qwen3-30B-A3B --seq 2048
loggetta train allenai/OLMoE-1B-7B-0924 --seq 512 --micro-batch 2 --steps 12 --out receipts/

Without --dataset, train runs a short demonstration on tatsu-lab/alpaca (full-sequence loss) with the default warmup + cosine learning-rate schedule. To train on your own data and keep a reusable adapter:

loggetta train Qwen/Qwen3-30B-A3B --dataset ./data/train.jsonl --format chat \
  --seq 1024 --micro-batch 1 --epochs 1 --out runs/my-run --adapter-out adapters/my-adapter

The dataset is validated, tokenized and profiled before any weights load, and the plan carries that profile. For chat and Alpaca data the defaults are assistant-only loss (--loss), isolated packing so each example attends only to itself (--packing), and warmup + cosine decay (--lr-schedule). Reload the adapter with loggetta.load_adapter("adapters/my-adapter"). See the training guide.

A feasible plan is an estimate, not an OOM guarantee. Serving placement can be planned; Loggetta does not launch the server yet.

Plan, run, measure

A plan is a serializable artifact, not just terminal output. It records the selected backend setup, budgets, memory estimates, rejected alternatives, warnings, and anything the planner does not model. You can save it, inspect it, then execute that exact decision:

loggetta plan allenai/OLMoE-1B-7B-0924 \
  --seq 512 --micro-batch 2 --steps 12 --out plan.json

loggetta execute plan.json --out receipts/

The resulting receipt records the plan plus measured allocator/reserved/driver memory, timing, correctness checks, and runtime provenance. Earlier receipts can be passed back as observations so later plans use measurements from the same model/setup when available:

loggetta plan allenai/OLMoE-1B-7B-0924 \
  --seq 512 --micro-batch 2 --observations receipts/

The planner does not load weights while choosing. With --dataset, every row is validated and tokenized first; then the model topology and the hardware are read. Candidate selection is deterministic policy over those facts, the data profile and the constraints.

Measured speed

These are measurements of the runtime/kernel stack that Loggetta installs and dispatches into, not speedups caused by the planner.

Result Measured comparison
2.352x faster/step vs Unsloth Qwen3-30B-A3B QLoRA on one RTX 5090, matched adapters/init/tokens and the same torch 2.12.1+cu130 / transformers 5.5.0 stack: 3.494 vs 8.218 s/step. Held-out loss was COMPARABLE. Unsloth used less peak VRAM: 24.27 vs 27.49 GB.
2.468x replication Same-stack Qwen3 comparison on a second RTX 5090 host. The registered position remains 2.352x; the replication is reported separately.
2.775x vs Axolotl Matched-work Qwen3-30B-A3B comparison on an RTX 5090 / Ryzen 9 9950X3D host: 2.147 vs 5.956 s/step. A separate EPYC-host reading was 1.979x, so this ratio is host-sensitive.
2.33x packed-compute throughput H100 synthetic expert-offload pipeline vs bitsandbytes CUDA dequantization + cuBLAS: 6.466 vs 2.773 pipeline tok/s, 26.8 vs 59.1 J/token. This is not an end-to-end serving claim.

Evidence: Unsloth same-stack · second-host replication · Axolotl/Unsloth matched work · H100 pipeline

Models: supported vs tested

Runtime-supported means the included experts4bit-qlora fast-training path is evidence-gated supported with a real-weight PASS receipt. The Loggetta column says what this repository itself has exercised.

Model / family Included runtime QLoRA Loggetta test status
OLMoE-1B-7B-0924 (olmoe) Supported Run: training + serving
Qwen3-30B-A3B (qwen3_moe) Supported Planner + serving validated; backend training receipts imported for calibration
Granite-3.1-3B-A800M (granitemoe) Supported Run: training + serving
Granite-4.0-H-tiny (granitemoehybrid) Supported Run: training; serving refusal validated (Mamba state unsupported by current paged runner)
Qwen3.6-35B-A3B (qwen3_5_moe) Supported Planner-tested for training; serving validated; training not yet executed here
LFM2-8B-A1B (lfm2_moe) Supported Planner-tested; serving refusal validated (conv state unsupported by current paged runner)
Mixtral-8x7B-Instruct-v0.1 (mixtral) Supported Planner-tested; not yet executed here
ERNIE-4.5-21B-A3B (ernie4_5_moe) Supported Planner-tested for training; serving validated (tiered, RTX A2000); training not yet executed here
Gemma-4-26B-A4B-it (gemma4_text) Supported Runtime-supported; not yet in Loggetta's model sweep
NVIDIA Nemotron-3.5-Lightning-30B-A3B (nemotron_h) Supported Runtime-supported; not yet in Loggetta's model sweep

Packing on hybrid models. With chat or Alpaca data the default isolated packing is refused for a model whose layers mix tokens through a state (Qwen3.6-35B-A3B, Granite-4.0-H-tiny, LFM2-8B-A1B and, by its structure, Nemotron-H) or whose attention is not described (DeepSeek-V2-Lite): resetting positions cannot isolate examples there. Pass --packing concat for those models.

Also exercised by the planner but not advertised as supported expert-QLoRA rows: DeepSeek-V2-Lite is planned with attention LoRA disabled because its MLA attention is not described by the current adapter path; gpt-oss-20b is deliberately refused for expert QLoRA because its biased/clamped expert structure does not satisfy ExpertsLoRA's contract.

Backend support is evidence-gated and can move independently of Loggetta's own test matrix. See the runtime capability register and Loggetta's measured results.

Memory planning

Loggetta labels each estimate as measured, derived, inferred, or heuristic instead of presenting every number as equally certain. An ExecutionReceipt records estimated vs measured memory and can be fed back into later plans.

Selected checks:

  • A2000 training runs across OLMoE and Granite families put allocator estimates within 0.01-0.21 GiB of measured peaks.
  • Qwen3-30B-A3B on an RTX 5090: allocator 22.09 GiB estimated vs 21.91 GiB measured; after receipt-calibrated runtime overheads, process peak 24.54 GiB planned vs 24.34 GiB measured.
  • OLMoE serving on an A2000: planned VRAM/DRAM/NVMe expert placement matched the server's 272 / 421 / 331 split; allocator 2.052 GiB planned vs 2.051 GiB measured.

Loggetta does not currently predict throughput. It may use measured evidence to order valid setups, but unknown speed remains unknown.

Current scope

Available: one-command install of the full stack; single-GPU MoE planning; saved ExecutionPlan and ExecutionReceipt; QLoRA training on your own data (local JSONL/JSON/CSV/Parquet/TXT or Hub datasets; text, Alpaca or chat) with reusable adapters; QLoRA training execution; device/host/storage serving placement planning; measured-memory feedback; explicit refusals.

Not claimed: first-class dense-model planning; multi-GPU planning/execution; calibrated throughput prediction; a Loggetta server-launch command; universal exposure of every mechanism in the lower packages.

GPU execution currently targets Linux + NVIDIA CUDA and requires a compatible driver/PyTorch environment. The package is pre-1.0.

More detail

The PyPI page is intentionally short. The GitHub repository carries the architecture, full benchmark caveats, receipts, family sweeps, and reproducibility notes:

Metadata

Release files for loggetta 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for loggetta 0.3.0
File Size Uploaded
loggetta-0.3.0.tar.gz 102.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for loggetta 0.3.0
File Interpreter ABI Platform
loggetta-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 175.0 kB

Release files / loggetta-0.3.0.tar.gz

Download URL loggetta-0.3.0.tar.gz
Size 102.3 kB
Tags Source
SHA-256 checksum
How to use checksums
2fbeb39c100f62f2507e6e055b265b0c24623c023c83320f760412a6c3c646c5
BLAKE2b-256 checksum
How to use checksums
b235b703c6c509331c300229ce3efc28ce470e62ab6d2f89f64911c6bfa0f35b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / loggetta-0.3.0-py3-none-any.whl

Download URL loggetta-0.3.0-py3-none-any.whl
Size 72.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9b9dad3d20d7e23e6e1c7f1a2fd55faa313b6f248c1eee05ed4e4ad7dcb2ba46
BLAKE2b-256 checksum
How to use checksums
537fbb7c2d58ea65ec1c1532d9f7551c9f0bcf6f6c8bde48d1689480375dc8ac
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

This release

0.3.0 This release

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page