Loggetta
One install for planning and running single-GPU MoE QLoRA workloads.
pip install loggetta
That installs Loggetta plus its current runtime and kernel stack:
- Loggetta — hardware inventory, planning,
ExecutionPlan, execution orchestration, andExecutionReceipt - experts4bit-qlora — model loading, QLoRA, training, serving mechanisms, and CPU/NVMe offload
- grouped-nf4-gemm — packed low-bit kernels and residency primitives
Loggetta examines the model, machine, workload, and constraints before loading model weights. It chooses a supported setup, explains why alternatives lost, and refuses configurations estimated not to fit. For supported training plans it dispatches into the included runtime and records what actually happened.
loggetta inspect
loggetta plan Qwen/Qwen3-30B-A3B --seq 2048
loggetta train allenai/OLMoE-1B-7B-0924 --seq 512 --micro-batch 2 --steps 12 --out receipts/
A feasible plan is an estimate, not an OOM guarantee. Serving placement can be planned; Loggetta does not launch the server yet.
Plan, run, measure
A plan is a serializable artifact, not just terminal output. It records the selected backend setup, budgets, memory estimates, rejected alternatives, warnings, and anything the planner does not model. You can save it, inspect it, then execute that exact decision:
loggetta plan allenai/OLMoE-1B-7B-0924 \
--seq 512 --micro-batch 2 --steps 12 --out plan.json
loggetta execute plan.json --out receipts/
The resulting receipt records the plan plus measured allocator/reserved/driver memory, timing, correctness checks, and runtime provenance. Earlier receipts can be passed back as observations so later plans use measurements from the same model/setup when available:
loggetta plan allenai/OLMoE-1B-7B-0924 \
--seq 512 --micro-batch 2 --observations receipts/
The planner does not load weights while choosing. Hardware probing and model topology discovery happen first; candidate selection is deterministic policy over those facts and constraints.
Measured speed
These are measurements of the runtime/kernel stack that Loggetta installs and dispatches into, not speedups caused by the planner.
| Result | Measured comparison |
|---|---|
| 2.352x faster/step vs Unsloth | Qwen3-30B-A3B QLoRA on one RTX 5090, matched adapters/init/tokens and the same torch 2.12.1+cu130 / transformers 5.5.0 stack: 3.494 vs 8.218 s/step. Held-out loss was COMPARABLE. Unsloth used less peak VRAM: 24.27 vs 27.49 GB. |
| 2.468x replication | Same-stack Qwen3 comparison on a second RTX 5090 host. The registered position remains 2.352x; the replication is reported separately. |
| 2.775x vs Axolotl | Matched-work Qwen3-30B-A3B comparison on an RTX 5090 / Ryzen 9 9950X3D host: 2.147 vs 5.956 s/step. A separate EPYC-host reading was 1.979x, so this ratio is host-sensitive. |
| 2.33x packed-compute throughput | H100 synthetic expert-offload pipeline vs bitsandbytes CUDA dequantization + cuBLAS: 6.466 vs 2.773 pipeline tok/s, 26.8 vs 59.1 J/token. This is not an end-to-end serving claim. |
Evidence: Unsloth same-stack · second-host replication · Axolotl/Unsloth matched work · H100 pipeline
Models: supported vs tested
Runtime-supported means the included experts4bit-qlora fast-training path is evidence-gated supported with a real-weight PASS receipt. The Loggetta column says what this repository itself has exercised.
| Model / family | Included runtime QLoRA | Loggetta test status |
|---|---|---|
OLMoE-1B-7B-0924 (olmoe) |
Supported | Run: training + serving |
Qwen3-30B-A3B (qwen3_moe) |
Supported | Planner + serving validated; backend training receipts imported for calibration |
Granite-3.1-3B-A800M (granitemoe) |
Supported | Run: training + serving |
Granite-4.0-H-tiny (granitemoehybrid) |
Supported | Run: training; serving refusal validated (Mamba state unsupported by current paged runner) |
Qwen3.6-35B-A3B (qwen3_5_moe) |
Supported | Planner-tested for training + serving; not yet executed here |
LFM2-8B-A1B (lfm2_moe) |
Supported | Planner-tested; serving refusal validated (conv state unsupported by current paged runner) |
Mixtral-8x7B-Instruct-v0.1 (mixtral) |
Supported | Planner-tested; not yet executed here |
ERNIE-4.5-21B-A3B (ernie4_5_moe) |
Supported | Planner-tested; not yet executed here |
Gemma-4-26B-A4B-it (gemma4_text) |
Supported | Runtime-supported; not yet in Loggetta's model sweep |
NVIDIA Nemotron-3.5-Lightning-30B-A3B (nemotron_h) |
Supported | Runtime-supported; not yet in Loggetta's model sweep |
Also exercised by the planner but not advertised as supported expert-QLoRA rows: DeepSeek-V2-Lite is planned with attention LoRA disabled because its MLA attention is not described by the current adapter path; gpt-oss-20b is deliberately refused for expert QLoRA because its biased/clamped expert structure does not satisfy ExpertsLoRA's contract.
Backend support is evidence-gated and can move independently of Loggetta's own test matrix. See the runtime capability register and Loggetta's measured results.
Memory planning
Loggetta labels each estimate as measured, derived, inferred, or heuristic instead of presenting every number as equally certain. An ExecutionReceipt records estimated vs measured memory and can be fed back into later plans.
Selected checks:
- A2000 training runs across OLMoE and Granite families put allocator estimates within 0.01-0.21 GiB of measured peaks.
- Qwen3-30B-A3B on an RTX 5090: allocator 22.09 GiB estimated vs 21.91 GiB measured; after receipt-calibrated runtime overheads, process peak 24.54 GiB planned vs 24.34 GiB measured.
- OLMoE serving on an A2000: planned VRAM/DRAM/NVMe expert placement matched the server's 272 / 421 / 331 split; allocator 2.052 GiB planned vs 2.051 GiB measured.
Loggetta does not currently predict throughput. It may use measured evidence to order valid setups, but unknown speed remains unknown.
Current scope
Available: one-command install of the full stack; single-GPU MoE planning; saved ExecutionPlan and ExecutionReceipt; QLoRA training execution; device/host/storage serving placement planning; measured-memory feedback; explicit refusals.
Not claimed: first-class dense-model planning; multi-GPU planning/execution; calibrated throughput prediction; a Loggetta server-launch command; universal exposure of every mechanism in the lower packages.
GPU execution currently targets Linux + NVIDIA CUDA and requires a compatible driver/PyTorch environment. The package is pre-1.0.
More detail
The PyPI page is intentionally short. The GitHub repository carries the architecture, full benchmark caveats, receipts, family sweeps, and reproducibility notes:
Metadata
Release files for loggetta 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| loggetta-0.1.3.tar.gz | 60.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| loggetta-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 107.1 kB
Release files / loggetta-0.1.3.tar.gz
| Download URL | loggetta-0.1.3.tar.gz |
|---|---|
| Size | 60.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e64a5e2791a050ae83d6461ec6a2747972db85eb044066bc1011547e085d3768
|
|
BLAKE2b-256 checksum How to use checksums |
c56090ecaab96f524837167a064e4fd894bf9090ba8edf2093e7c13977e02e79
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / loggetta-0.1.3-py3-none-any.whl
| Download URL | loggetta-0.1.3-py3-none-any.whl |
|---|---|
| Size | 46.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d41df97c78f5322893df13778bb296367a73c10759282144388317ea08102988
|
|
BLAKE2b-256 checksum How to use checksums |
ddcd827ea86453ffa26925a648780a70659b6106e172fc9c9367834d0ec9a0b1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log