Skip to main content

Loggetta

One install. Plan. Run. Measure.

The user-facing home for experts4bit-qlora and grouped-nf4-gemm.

Plan a workload for your machine, run supported training through the included runtime, and keep the receipt.

PyPI Python >=3.10 pre-1.0 single-GPU MoE ExecutionPlan + ExecutionReceipt

pip install loggetta

That installs the stack: Loggetta, the experts4bit-qlora runtime with its training and fast-kernel dependencies, and grouped-nf4-gemm. No backend extra or separate package assembly is required.

Measured speed

The performance work lives in the runtime and kernel layers that Loggetta installs and dispatches into. These are measured positions for that included stack, not speedups produced by the planner itself.

Result Measured comparison
2.352× faster per training step vs Unsloth Qwen3-30B-A3B QLoRA on an RTX 5090, matched adapters, initialization and tokens, with both frameworks on torch 2.12.1+cu130 / transformers 5.5.0: 3.494 vs 8.218 s/step. Held-out loss at N=60 was 0.7569 vs 0.7557 and graded COMPARABLE. Unsloth used less peak VRAM: 24.27 vs 27.49 GB.
2.468× replication The same-stack Qwen3 comparison on a second RTX 5090 host: e4b 4.143 / 4.073 s/step vs Unsloth 10.155 / 10.120 s/step. The registered position remains 2.352×; the replication is quoted separately.
2.775× vs Axolotl Matched-work Qwen3-30B-A3B comparison on an RTX 5090 / Ryzen 9 9950X3D host: 2.147 vs 5.956 s/step. A separate EPYC-host reading was 1.979×, so this ratio is explicitly host-sensitive.
2.33× packed-compute throughput H100 synthetic expert-offload pipeline, grouped-nf4-gemm vs bitsandbytes CUDA dequantization + cuBLAS: 6.466 vs 2.773 pipeline tok/s, 26.8 vs 59.1 J/token. This is a pipeline result, not an end-to-end serving claim.

Evidence: same-stack Unsloth result · second-host replication · matched Axolotl/Unsloth result · H100 bnb comparison

loggetta inspect
loggetta plan Qwen/Qwen3-30B-A3B --seq 2048

Loggetta turns a workload, a machine, constraints, and evidence into an ExecutionPlan. It explains the selected configuration, records why alternatives lost, and refuses unsupported or infeasible plans before loading model weights. For supported training plans, Loggetta dispatches into e4b and records the result as an ExecutionReceipt. Serving placement planning is also included; starting the server remains a separate e4b step in this release.

Planner ExecutionPlan ExecutionReceipt
Turns hardware, workload, constraints, objectives, and evidence into a decision. Captures the selected backend/setup, budgets, estimates, rejected alternatives, refusal reasons, warnings, and evidence quality. Records what actually happened: memory, timing, correctness, engagement, provenance, and estimate-vs-measured deltas.

Quick start

After installation, use the loggetta command throughout:

# Inspect the machine
loggetta inspect

# Choose a supported configuration before loading weights
loggetta plan Qwen/Qwen3-30B-A3B --seq 2048

# Constrain the planner deliberately: device budget is in GiB
loggetta plan allenai/OLMoE-1B-7B-0924 --experts device --vram 6

# Save a training plan
loggetta plan allenai/OLMoE-1B-7B-0924 \
  --seq 512 --micro-batch 2 --steps 12 --out plan.json

# Execute the saved plan through the included runtime and write its receipt
loggetta execute plan.json --out receipts/

# Plan again using measured evidence from earlier runs
loggetta plan allenai/OLMoE-1B-7B-0924 \
  --seq 512 --micro-batch 2 --observations receipts/

For a short training run without a separate save/execute step:

loggetta train allenai/OLMoE-1B-7B-0924 \
  --seq 512 --micro-batch 2 --steps 12 --out receipts/

train combines planning and execution; it does not bypass admission checks. The equivalent python -m loggetta ... commands are also supported.

For serving placement:

loggetta plan Qwen/Qwen3-30B-A3B \
  --workload serve --context 4096 --concurrency 1

The serving plan's Why carries the server environment. loggetta execute does not start servers yet. The runtime and kernels are installed, but not every lower-layer capability has a Loggetta execution command.

Three levels of control

Mode Example What it means
Automatic loggetta plan MODEL Let the planner choose among supported candidates.
Directed --vram, --ram, --experts, --objective State the budget or goal without hand-building a backend setup.
Expert --fix FIELD=VALUE Pin a backend setup field and make the remaining search work around it.

How the stack runs

The CLI/API, planner, plan, execution handoff, and receipt belong to Loggetta. Execution calls the included runtime; the runtime calls its kernels. Backend packages remain separate projects without becoming separate setup chores.

flowchart TB
    U["User: pip install loggetta"] --> P
    I["Workload + machine<br/>constraints + evidence"] --> P
    subgraph L["Loggetta: CLI / API and orchestration"]
        P["Planner"] --> X["ExecutionPlan<br/>what should run + why"]
        X --> O["execute(plan)<br/>validate + dispatch"]
        R["ExecutionReceipt<br/>what actually happened"] -. "measured feedback" .-> P
    end
    O --> E["experts4bit-qlora<br/>runs the selected backend setup"]
    E --> G["grouped-nf4-gemm<br/>kernels + residency primitives"]
    E -. "execution results" .-> R

Installed together. Responsibilities kept separate.

Layer Owns Does not reimplement
Loggetta The default stack install and user-facing CLI/API; hardware inventory/provenance; budgets, objectives, constraints, candidate ordering, refusal policy, ExecutionPlan, ExecutionReceipt, execution orchestration Model loaders, QLoRA machinery, serving engines, residency engines, or kernels
experts4bit-qlora Model topology/conventions, admission, footprint primitives tied to its runtime, loading, adapters, QLoRA preparation, training, serving, residency/offload mechanisms Loggetta's planning and orchestration policy
grouped-nf4-gemm Packed low-bit kernels, routing/capability facts, and low-level residency primitives Model-level execution or workload planning
Loggetta -> experts4bit-qlora -> grouped-nf4-gemm

No cycles. No duplicated topology. No plugin framework until a second backend actually earns one.

plan() is pure policy: for the same inputs it returns the same serializable ExecutionPlan, without loading weights or touching the device. Hardware probing and model-description gathering happen before that pure planning step. The plan includes the winner, every relevant loser and why it lost, budget sources, evidence labels, warnings, and explicit "not modeled" items.

execute(plan) validates the plan, dispatches through the selected backend executor, and writes the resulting ExecutionReceipt. Before anything loads, it refuses a refused plan, a workload kind its backend only plans, and a plan made for a different GPU (plan --hardware), whose receipt would name the wrong card. The backend receives its own setup; it does not need to import Loggetta's plan types.

A real plan, abridged (OLMoE-1B-7B on an RTX A2000, from evidence/2026-10-04-rtx-a2000/)
Budget    device  10.65 GiB [reported: free now]   host  23.49 GiB [reported: available now]   headroom   0.53 GiB [policy]

Selected  backend experts4bit: experts on device, grouped_nf4 kernel

Estimated memory (each line says how it is known)
  device frozen expert stacks                    3.38 GiB  [derived]  16 stacks in nf4, blocksize 64 (packed + absmax)
  device dense weights (bf16)                    0.89 GiB  [derived]
  device optimizer state (adamw)                 0.23 GiB  [derived]
  device activations                             0.54 GiB  [heuristic]  16 saved layer inputs (T x H bf16) + ...
  device allocator reserve (cached, unallocated blocks)   1.17 GiB  [measured]  receipt ...173932Z: ... = 0.222
  device CUDA context + library workspaces       0.13 GiB  [measured]  receipt ...173932Z
  device total                                   6.56 GiB   of  10.65 GiB budget

Why
  - experts resident on the device: 6.56 GiB estimated + 0.53 GiB headroom fits the 10.65 GiB budget
  - ordering, expert_kernel: grouped_nf4 before reference: e4b.train.h2h.unsloth.olmoe.5090.2026-09-19 (1.39 vs 14.88 s/step), ...

Performance  not predicted: no performance model is calibrated for this backend yet; ...

Alternatives considered
  [fits ] experts on device, grouped_nf4 kernel, NF4 attention       device   6.10 GiB  host   0.72 GiB  ...
  [fits ] experts on host, grouped_nf4 kernel                        device   2.69 GiB  host   4.10 GiB  ...
  ...

Its receipt: allocator 5.26 GiB estimated, 5.47 GiB measured; driver 6.56 GiB estimated, 6.82 GiB measured.

Evidence, not vibes

Loggetta is deliberately not a performance oracle. It records what it knows, labels heuristics, refuses plans it cannot support, and feeds measured receipts back into future planning.

Current checks from docs/RESULTS.md:

Check Measured result
Planner parity, A2000 / OLMoE Planned and hand-composed paths had bitwise-identical step-1 loss (1.858969), effectively identical allocator peaks (5.4710 vs 5.4706 GiB), and the same 60,817,408 trainable parameters.
Allocator estimates, A2000 Across six measured runs spanning three model families, allocator estimates were within 0.01 to 0.21 GiB of measured peaks.
Qwen3-30B-A3B, RTX 5090 Allocator estimate 22.09 GiB, measured 21.91 GiB. After receipt-calibrated runtime overheads: planned process peak 24.54 GiB, measured 24.34 GiB.
Serving tiers, A2000 / OLMoE The planned VRAM / DRAM / NVMe split matched the server's own (272 / 421 / 331 experts); allocator 2.052 GiB planned, 2.051 GiB measured.
Host offload Transfer time is treated as a lower bound, not a fabricated step-time prediction.
Included runtime speed, Qwen3-30B-A3B / RTX 5090 Same-stack matched-work training measured 3.494 s/step for e4b vs 8.218 s/step for Unsloth (2.352×), with comparable held-out loss; replicated at 2.468× on a second host.

An ExecutionReceipt carries provenance, allocator / reserved / driver peaks, correctness checks, engagement evidence, timing, and estimate-vs-measured deltas. That receipt is the measured counterpart to the plan that produced it.

Current scope

Available today Not claimed yet
Default installation of Loggetta, e4b, and gnf4 First-class dense-model planning
Single-GPU MoE planning Multi-GPU placement or execution planning
Saved ExecutionPlan and ExecutionReceipt artifacts Calibrated throughput prediction
QLoRA execution through loggetta execute and loggetta train Profile-driven serving hot sets
Serving placement planning across device, host, and storage tiers A Loggetta server-launch command
Measured-memory feedback, explicit constraints, and inspectable refusals Universal support for every mechanism exposed by the lower packages

Those are roadmap items, not assumptions hidden inside current results.

Installation details

Starting with 0.1.1, the default install includes the runtime and kernel packages. There is no planner-only default that needs an extra to become useful.

Installed package Role
Loggetta Planner, ExecutionPlan, ExecutionReceipt, CLI/API, orchestration, and measured feedback
experts4bit-qlora[train,fast] >= 0.48.0 Model loading, QLoRA, training, serving, and residency, with the training and fast-path dependencies enabled
grouped-nf4-gemm >= 0.41.0 Packed low-bit kernels and residency primitives
datasets Data loading used by the training path

Upgrading an earlier installation:

python -m pip install --upgrade loggetta

The old loggetta[experts4bit] spelling remains accepted as a compatibility alias. It is no longer necessary. Both lower packages can still be installed and used independently.

The minimum backend releases support training plans and execution plus serving plans. The newer serving-estimate refinements in RESULTS 6b-6e require APIs beyond e4b 0.48.0; that minimum release is not sufficient to reproduce those specific measurements. See docs/RESULTS.md for the backend revisions used.

Source installation for development:

git clone https://github.com/pjordanandrsn/loggetta
cd loggetta
python -m pip install -e .

Why keep the packages separate?

One install does not require one codebase. Kernel selection belongs with kernels. Model topology and QLoRA mechanisms belong with the runtime. Loggetta brings them together, decides which combination should run on this machine under the user's constraints, and makes that decision inspectable before execution.

The separation lets e4b and gnf4 remain independently useful, keeps backend knowledge with its owner, and allows receipts to improve planning policy without moving model or kernel code into the planner. Users get one starting point; maintainers keep clear boundaries.

Documentation

Read this For
SESSION-REPORT.md Short answers and current state
ARCHITECTURE.md Ownership boundaries and the Planner -> Plan -> Backend -> Receipt model
RESULTS.md Measurements, estimator error, receipts, and what changed because of them
SERVING-PRESSURE-TEST.md Serving placement and pressure-test evidence
Tests and validation commands
python -m pip install -e ".[test]"
pytest tests/
python bench/direct_baseline.py --receipt R.json
python bench/validate_register.py
python bench/family_sweep.py --hardware hw.json
python bench/summarize_receipts.py runs/receipts

The fast test suite is CPU-only. Planner tests use the installed backend packages; individual tests may skip when a platform cannot import a required kernel module. Execution handoff tests use a stand-in executor rather than launching training. GPU benchmarks and validation runs remain separate.


Install Loggetta. Plan. Run. Measure. Feed the receipt back.

Metadata

Release files for loggetta 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for loggetta 0.1.2
File Size Uploaded
loggetta-0.1.2.tar.gz 61.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for loggetta 0.1.2
File Interpreter ABI Platform
loggetta-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 112.1 kB

Release files / loggetta-0.1.2.tar.gz

Download URL loggetta-0.1.2.tar.gz
Size 61.9 kB
Tags Source
SHA-256 checksum
How to use checksums
1a7d5a4e293be09a6efc8b5a8bbbe58b25d960e1eaf2269498fbb5ae4e434cb7
BLAKE2b-256 checksum
How to use checksums
02e7373253e836d2231ba71b0d24e82b21f6ad987521899f38e2077143c795d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / loggetta-0.1.2-py3-none-any.whl

Download URL loggetta-0.1.2-py3-none-any.whl
Size 50.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a753356bc5af97a443b6b2e9a0a5192d9bfa571ba881b15f6b89db05148db9ca
BLAKE2b-256 checksum
How to use checksums
eee005693c7ccf084106a5c1060a8da65fa71ba818e49a96eeb59464e7930d9e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.1.3

2 release files

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page