Skip to main content

Loggetta

Plan first. Execute through the backend.

Loggetta turns a workload, a machine, and constraints into an inspectable execution plan.

It selects a supported configuration, explains why it won, records why alternatives lost, and refuses impossible plans before loading weights.

Python >=3.10 PyPI pre-1.0 single-GPU MoE ExecutionPlan + ExecutionReceipt


Loggetta owns the Planner, the ExecutionPlan, and the ExecutionReceipt. Its current focus is single-GPU MoE training and serving through experts4bit-qlora, with grouped-nf4-gemm providing the packed low-bit kernel and residency primitives beneath that runtime.

Planner ExecutionPlan ExecutionReceipt
Turns hardware, workload, constraints, objectives, and evidence into a decision. Captures the selected backend/setup, budgets, estimates, rejected alternatives, refusal reasons, warnings, and evidence quality. Records what actually happened: memory, timing, correctness, engagement, provenance, and estimate-vs-measured deltas.

The loop

flowchart LR
    I["Workload + machine<br/>constraints + evidence"] --> P["Loggetta Planner"]
    P --> X["ExecutionPlan<br/>what should run + why"]
    X --> E["experts4bit-qlora<br/>executes the plan"]
    E --> G["grouped-nf4-gemm<br/>kernels + primitives"]
    E --> R["ExecutionReceipt<br/>what actually happened"]
    R -. "measured feedback" .-> P

The ownership split is deliberate:

Layer Owns Does not own
Loggetta Hardware inventory/provenance, budgets, objectives, constraints, candidate ordering, refusal policy, ExecutionPlan, ExecutionReceipt, thin execution orchestration Model loaders, QLoRA machinery, serving engines, residency engines, kernels
experts4bit-qlora Model topology/conventions, admission, footprint primitives tied to its runtime, loading, adapters, QLoRA preparation, training, serving, residency/offload mechanisms Cross-backend planning policy
grouped-nf4-gemm Packed low-bit kernels, routing/capability facts, and low-level residency primitives Model/runtime policy or planner decisions
Loggetta -> experts4bit-qlora -> grouped-nf4-gemm

No cycles. No duplicated topology. No plugin framework until a second backend actually earns one.

Quick start

# What machine am I actually on?
python -m loggetta inspect

# Produce the artifact Loggetta exists to produce
python -m loggetta plan Qwen/Qwen3-30B-A3B --seq 2048

# Constrain the planner deliberately
python -m loggetta plan allenai/OLMoE-1B-7B-0924 --experts device --vram 6

# Keep the plan as a file, then execute that plan through the backend; the receipt lands in receipts/
python -m loggetta plan allenai/OLMoE-1B-7B-0924 --seq 512 --micro-batch 2 --steps 12 --out plan.json
python -m loggetta execute plan.json --out receipts/

# Plan again from what was measured
python -m loggetta plan allenai/OLMoE-1B-7B-0924 --seq 512 --micro-batch 2 --observations receipts/

train MODEL ... is plan and execute in one step. Serving is planned with plan MODEL --workload serve --context T --concurrency N. The plan's Why carries the server's exact environment; execute does not start servers yet.

Three levels of control

Mode Example What it means
Automatic plan MODEL Let the planner choose among supported candidates.
Directed --vram, --ram, --experts, --objective State the budget or goal without hand-building a backend setup.
Expert --fix FIELD=VALUE Pin a backend setup field and make the remaining search work around it.

plan() is pure policy: for the same inputs it returns the same serializable ExecutionPlan, without loading weights or touching the device. The plan includes the winner, every relevant loser and why it lost, budget sources, evidence labels, warnings, and explicit “not modeled” items.

execute(plan) is intentionally thin. It validates the plan, delegates execution to the selected backend, and constructs an ExecutionReceipt from the backend result. It does not reimplement the runtime. Before anything loads, it refuses a refused plan, a workload kind its backend only plans, and a plan made for a different GPU (plan --hardware), whose receipt would name the wrong card.

A real plan, abridged (OLMoE-1B-7B on an RTX A2000, from evidence/2026-10-04-rtx-a2000/)
Budget    device  10.65 GiB [reported: free now]   host  23.49 GiB [reported: available now]   headroom   0.53 GiB [policy]

Selected  backend experts4bit: experts on device, grouped_nf4 kernel

Estimated memory (each line says how it is known)
  device frozen expert stacks                    3.38 GiB  [derived]  16 stacks in nf4, blocksize 64 (packed + absmax)
  device dense weights (bf16)                    0.89 GiB  [derived]
  device optimizer state (adamw)                 0.23 GiB  [derived]
  device activations                             0.54 GiB  [heuristic]  16 saved layer inputs (T x H bf16) + ...
  device allocator reserve (cached, unallocated blocks)   1.17 GiB  [measured]  receipt ...173932Z: ... = 0.222
  device CUDA context + library workspaces       0.13 GiB  [measured]  receipt ...173932Z
  device total                                   6.56 GiB   of  10.65 GiB budget

Why
  - experts resident on the device: 6.56 GiB estimated + 0.53 GiB headroom fits the 10.65 GiB budget
  - ordering, expert_kernel: grouped_nf4 before reference: e4b.train.h2h.unsloth.olmoe.5090.2026-09-19 (1.39 vs 14.88 s/step), ...

Performance  not predicted: no performance model is calibrated for this backend yet; ...

Alternatives considered
  [fits ] experts on device, grouped_nf4 kernel, NF4 attention       device   6.10 GiB  host   0.72 GiB  ...
  [fits ] experts on host, grouped_nf4 kernel                        device   2.69 GiB  host   4.10 GiB  ...
  ...

Its receipt: allocator 5.26 GiB estimated, 5.47 GiB measured; driver 6.56 GiB estimated, 6.82 GiB measured.

Evidence, not vibes

Loggetta is deliberately not a performance oracle. It records what it knows, labels heuristics, refuses plans it cannot support, and feeds measured receipts back into future planning.

Current checks from docs/RESULTS.md:

Check Measured result
Planner parity, A2000 / OLMoE Planned and hand-composed paths had bitwise-identical step-1 loss (1.858969), effectively identical allocator peaks (5.4710 vs 5.4706 GiB), and the same 60,817,408 trainable parameters.
Allocator estimates, A2000 Across six measured runs spanning three model families, estimates landed +0.01 to +0.21 GiB from measured peaks.
Qwen3-30B-A3B, RTX 5090 Allocator estimate 22.09 GiB, measured 21.91 GiB. After receipt-calibrated runtime overheads: planned process peak 24.54 GiB, measured 24.34 GiB.
Serving tiers, A2000 / OLMoE The planned VRAM / DRAM / NVMe split matched the server's own (272 / 421 / 331 experts); allocator 2.052 GiB planned, 2.051 GiB measured.
Host offload Transfer time is treated as a lower bound, not a fabricated step-time prediction.

An ExecutionReceipt carries provenance, allocator / reserved / driver peaks, correctness checks, engagement evidence, timing, and estimate-vs-measured deltas. That receipt is the measured counterpart to the plan that produced it.

Current scope

Supported today Not claimed yet
  • ✅ Single-GPU MoE planning
  • ✅ First-class ExecutionPlan and ExecutionReceipt artifacts
  • ✅ QLoRA planning with execution delegated to experts4bit-qlora
  • ✅ Serving placement planning across device, host, and storage tiers
  • ✅ Measured-memory feedback through receipts
  • ✅ Deterministic, inspectable plans and refusals
  • ✅ Explicit constraints instead of hidden “magic” defaults
  • ⏳ Dense-model planning as a first-class Loggetta backend
  • ⏳ Multi-GPU placement or execution planning
  • ⏳ Calibrated throughput prediction
  • ⏳ Profile-driven serving hot sets
  • ⏳ Automatic execution of every serving plan

Those are roadmap items, not assumptions hidden inside current results.

Install

Loggetta is the user-facing install for the stack. One command installs the planner, the e4b runtime, and the gnf4 kernel layer:

python -m pip install loggetta

That installs:

  • Loggetta — Planner, ExecutionPlan, ExecutionReceipt, CLI, orchestration and measured feedback
  • experts4bit-qlora >= 0.48.0 — model/runtime layer for loading, QLoRA, training, serving and residency
  • grouped-nf4-gemm >= 0.41.0 — packed low-bit kernels and residency primitives

The old loggetta[experts4bit] spelling remains accepted as a compatibility alias, but the extra is no longer required.

The released backend versions support training plans and execution plus serving plans. The newest serving-estimate refinements in RESULTS 6b–6e depend on APIs newer than e4b 0.48.0; reproducing those specific measurements currently requires experts4bit-qlora main. Loggetta labels unavailable mechanisms rather than pretending they exist.

Source installs remain useful for development:

git clone https://github.com/pjordanandrsn/loggetta
cd loggetta
python -m pip install -e .

Why a separate planner?

Kernel selection belongs with kernels. Model topology and QLoRA mechanisms belong with the runtime. The decision which combination should run on this machine, under this budget, for this objective is a separate concern, and the ExecutionPlan is the durable artifact of that decision.

Keeping that boundary clean means:

  • backend packages remain independently useful;
  • planning stays deterministic and inspectable;
  • the plan can explain both selection and refusal before weight loading;
  • the runtime can evolve without absorbing cross-machine policy;
  • receipts can improve future policy without contaminating kernel or model code;
  • a second backend can earn a general abstraction instead of forcing one prematurely.

Documentation

Read this For
SESSION-REPORT.md Short answers and current state
ARCHITECTURE.md Ownership boundaries and the Planner → Plan → Backend → Receipt model
RESULTS.md Measurements, estimator error, receipts, and what changed because of them
SERVING-PRESSURE-TEST.md Serving placement and pressure-test evidence
Tests and validation commands
pytest tests/
python bench/direct_baseline.py --receipt R.json
python bench/validate_register.py
python bench/family_sweep.py --hardware hw.json
python bench/summarize_receipts.py runs/receipts

The fast test suite is CPU-only. Planner tests run when experts4bit-qlora is importable. The execution tests check the handoff to the backend with a stand-in executor, so they need nothing installed. GPU benchmarks and validation runs stay separate so ordinary correctness tests do not silently become hardware-dependent.


Plan first. Execute through the backend. Measure. Feed the receipt back.

Metadata

Release files for loggetta 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for loggetta 0.1.1
File Size Uploaded
loggetta-0.1.1.tar.gz 58.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for loggetta 0.1.1
File Interpreter ABI Platform
loggetta-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 106.2 kB

Release files / loggetta-0.1.1.tar.gz

Download URL loggetta-0.1.1.tar.gz
Size 58.0 kB
Tags Source
SHA-256 checksum
How to use checksums
dffc0b20fd50e6b2c2e7a5d35bd52db8db36ec12e4101e6f5f33b9a9815719b3
BLAKE2b-256 checksum
How to use checksums
3dd7d840b0657ff88f338b3b3bf27fb657043e88b1cf06d41f599a9a051496ed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / loggetta-0.1.1-py3-none-any.whl

Download URL loggetta-0.1.1-py3-none-any.whl
Size 48.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ac6ef442fdecff4379f37cd9bd2556623761cc299b9a671cc302b0a1e135fc1
BLAKE2b-256 checksum
How to use checksums
0e3fe12872983dfeb720670d88f66fbf9c51fd2ae427fda0971028c5934db642
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page