Skip to main content

tinyforge

Train, evaluate and serve small LLMs from scratch on modest GPUs (developed against a 4 GB RTX 3050 Ti). One engine, three front doors: CLI (for developers/CI), REST API, and a web UI (which can later be wrapped as a desktop app with Tauri/Electron without changing the backend).

Status: early alpha (0.1). Tested on Windows with one NVIDIA GPU; Linux passes the test suite in Docker but has not been run on a real GPU box. Install: pip install tinyforge, then tinyforge doctor to see what else you need (PyTorch, pip install "tinyforge[finetune]", and optional pieces such as Soup and llama.cpp).

Quick start

pip install torch --index-url https://download.pytorch.org/whl/cu124   # or /cpu
pip install -e ".[dev]"            # add ,triton for fused kernels
tinyforge doctor                   # hardware + warnings, with fixes
tinyforge pipeline --preset micro --steps 2000   # doctor -> data -> plan -> train -> eval
tinyforge generate "ROMEO:" --int8
tinyforge serve                    # UI at http://127.0.0.1:8000

What you get

Stage What it does Guard rails (warning codes)
doctor GPU/VRAM/bf16/Triton/JAX probe HW001-HW008
data prepare download or ingest text, train BPE, contiguous train/val split DQ001-DQ005: tiny corpus, duplicates, bad vocab
plan VRAM estimate; auto-enables grad checkpointing, shrinks micro-batch, keeps effective batch PL001-PL005
train AMP (bf16/fp16+scaler), grad accumulation, cosine LR, atomic checkpoints, auto-resume TR001 NaN skip/abort, TR002 overfit, TR003 plateau
eval perplexity vs random baseline, diversity, memorisation, determinism, KV-cache correctness, speed/VRAM EV001-EV009, non-zero exit on failure
generate/serve KV-cache decoding, top-k/top-p, int8 weights, optional Triton RMSNorm

Fine-tuning platform commands

Command What it does
memory probe VRAM/RAM/disk and say where a model's frozen weights would live and the estimated speed
data ingest documents (txt, md, html; pdf/docx/pptx/xlsx with the docs extra) -> cleaned, chunked, de-duplicated JSONL
data tabular CSV -> text-to-SQL examples whose answers are verified by executing them (--hard for eval shapes)
data pairs chunks -> grounded Q&A pairs from a teacher model (any OpenAI-compatible endpoint), split before generation
data pii-scan, --pii find emails, phones, cards, IDs and credentials; flag, redact or drop
worker run --spec job.yaml run a TrainingJob; backend: native or soup (layer streaming for models larger than VRAM)
export gguf adapter + base -> GGUF for llama.cpp
bench list/suggest/run pick benchmarks sized to your GPU/RAM/time budget (MMLU, ARC, HellaSwag, GSM8K, IFEval via lm-evaluation-harness, plus a built-in SQL execution check) and run them, base vs tuned
serve [--engine llamacpp] API + UI; OpenAI-compatible /v1/chat/completions with guardrails and metrics

Verified on one RTX 3050 Ti 4 GB laptop (Windows): a 3B and an 8B model fine-tuned through the job spec and served locally; see benchmarks/ and docs/soup-backend.md. Early alpha: single GPU, Windows tested only.

Model: decoder-only transformer with RoPE, RMSNorm, SwiGLU, tied embeddings, PyTorch SDPA (FlashAttention kernels). Presets: nano ~1M, micro ~12M, small ~28M, base ~100M.

Fine-tune a pretrained model (LoRA / QLoRA)

The most useful thing a small GPU can do. Default: SmolLM2-360M-Instruct on an instruction dataset.

pip install -e ".[finetune]"
tinyforge ft pipeline --steps 150          # doctor -> data -> plan -> train -> eval + merge
tinyforge ft generate "Explain hash tables" # tuned model;  add --base to compare with the original
Stage What it does Guard rails
ft data JSONL or HF dataset -> chat JSONL; dedupe; hash split keyed on prompt (no leakage); length stats FD001-FD006: too few examples, duplicates, malformed, tiny val set, truncation, PII/credentials (credentials are dropped)
ft plan picks fp16 vs 4-bit NF4, checkpointing, and a token budget per micro-batch to fit VRAM FP001-FP007
ft train LoRA, length-grouped token-budget batching, resume, best-adapter tracking FT001 NaN, FT002 overfit, FT005 VRAM spill, FT006 adaptive OOM recovery
ft eval tuned vs base held-out loss, forgetting check on general text, side-by-side generations, merge + equivalence check FE001-FE009; exit 1 if tuned is not better than base

Measured on an RTX 3050 Ti Laptop (4 GB), SmolLM2-360M, 150 steps x 16 examples: 2.4 GB peak VRAM, ~830 tok/s, ~10 min, held-out loss 1.314 -> 1.262 (-4.0%), no forgetting; exports a 690 MB standalone merged model. Lessons baked into the planner (all measured, see finetune.plan): small micro-batches are launch-bound (bs8 is 2.3x the tokens/s of bs4), LoRA dropout 0 is ~40% faster, and allocator-level estimates miss cuBLAS workspace, so training recovers from memory errors by shrinking the token budget instead of crashing.

Layout

src/tinyforge/   engine (config, data, model, train, evaluate, infer, kernels/), cli.py, server.py, ui/
tests/           unit + end-to-end smoke (synthetic data, no network)
infra/terraform/ S3 artifacts, IAM, no-inbound GPU spot worker (SSM access), budget alarm
.github/         CI: lint, test matrix, terraform validate, docker build, self-hosted GPU e2e
Dockerfile       CUDA runtime image

Cloud (AWS)

cd infra/terraform
terraform init -backend-config="bucket=<state-bucket>" -backend-config="key=tinyforge/tf.tfstate"
terraform apply -var enable_gpu_worker=true -var alert_email=you@example.com
# then run the `connect` output command to port-forward the UI over SSM (no open ports)

Verified results (what was actually run)

Everything below was run on one machine: Windows 11, RTX 3050 Ti Laptop GPU (4 GB), 16 GB RAM. Each row links to the raw numbers. These are single runs on small test sets, not benchmarks: read the caveats in each JSON file.

What was tested Result Data
Llama-3.1-8B fine-tuned on a 4 GB GPU (layer streaming through the job spec), served with llama.cpp trained 200 steps in 9.5 min; on 100 questions about a table it never saw: 100% execution accuracy vs 93% for the same model told to reply with SQL only e8-penguins-8b-soup.json
Same 8B model, harder question shapes never used in training 98% vs 97%: no meaningful gain over a well-prompted base model. Fine-tuning mostly fixed output format e8-penguins-8b-hard.json
Qwen2.5-3B fine-tuned end to end (data -> train -> evaluate -> serve) 98% execution accuracy vs 70% for the base model told to reply with SQL (0% without that instruction) e2e-titanic-3b-soup.json
Soup-trained adapter, quality check (3B) exact match 4% -> 45%, SQL that runs 9% -> 88% on 100 held-out rows soup-adapter-check-qwen3b.json
Chunked cross-entropy (memory saving), Qwen2.5-1.5B 4-bit peak GPU memory 3.65 -> 2.95 GB, 289 -> 499 tok/s, same loss chunked-ce-qwen1p5b-4bit-rerun.json
llama.cpp inference options, 3B flash attention ~8% faster; 8-bit KV cache halves the cache (144 -> 76.5 MB); n-gram speculation 129 vs 54.5 tok/s only on repeated requests (no gain on a first request) inference-techniques-qwen3b.json
A deliberately bad model (random weights, real architecture) the tool now fails it: FE010 base model looks untrained, FE011 degenerate looping output (before the fix it printed "passed"). A model without a chat template gets a clear FD007 error tests/test_bad_models.py
Test suite 290 tests pass on Windows (Python 3.12) and on Linux in Docker (Python 3.10 and 3.12, CPU PyTorch) pytest -q

What these results do not show: general-purpose quality gains (the SQL tests use templated questions), results on other hardware or operating systems, multi-GPU training, or the standard benchmark runner (bench run, built on lm-evaluation-harness), which has not been run end to end yet.

The dashboard serving a 3B model through llama.cpp: time to first token, throughput, latency, per-token latency, prefill speed, KV-cache growth and GPU memory per request

The dashboard while serving a 3B model. KV-cache use grows linearly with context (36 KiB per token) while GPU memory stays flat, because llama.cpp reserves the cache up front.

Honest limits / roadmap

  • JAX: there is a probe and a backend slot, but training is PyTorch-only today. JAX has no native Windows GPU support (use WSL2/Linux), and a second backend should be added behind the same train() interface.
  • Triton: one fused kernel (RMSNorm forward, inference only). Next: fused SwiGLU, then a backward pass. Triton on Windows needs the triton-windows package.
  • From-scratch models are for learning/prototyping; use fine-tuning for anything useful. A 4% held-out gain on a general instruction set is modest by design: LoRA on a small, already-tuned model mostly shifts style. Use a task-specific dataset to see larger gains.
  • 4-bit QLoRA is verified on real hardware (see ROADMAP.md): it works but ends ~4.7% worse in loss than fp16 on a model that fits in fp16. Its value is fitting larger models: an 8B model trains on a 4 GB GPU via backend: soup.
  • Generation speed (HF generate, eager) is ~10 tok/s on this GPU; the merged model is the thing to export to llama.cpp/vLLM for serving.
  • Single-GPU, single-job (a SQLite job queue exists). Multi-GPU (FSDP/DDP) is not built and untested.
  • The API requires TINYFORGE_API_TOKEN (bearer) and the server binds to localhost by default; the static UI page itself is not authenticated. Put it behind TLS, SSM/VPN or a reverse proxy before exposing it.

Metadata

Release files for tinyforge 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tinyforge 0.1.0
File Size Uploaded
tinyforge-0.1.0.tar.gz 180.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tinyforge 0.1.0
File Interpreter ABI Platform
tinyforge-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 328.7 kB

Release files / tinyforge-0.1.0.tar.gz

Download URL tinyforge-0.1.0.tar.gz
Size 180.6 kB
Tags Source
SHA-256 checksum
How to use checksums
0389b1614fbcc0f63f169883ed96469d2847d809042a0b1cf5537fd438db7403
BLAKE2b-256 checksum
How to use checksums
58897b163909e32a851ab919d92e1dc2d4a6db97631aed332e39b8376e6aca4e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / tinyforge-0.1.0-py3-none-any.whl

Download URL tinyforge-0.1.0-py3-none-any.whl
Size 148.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5f8781e05f19544b1e96cfe40357fb0d11162c99f4cca66bd2554a7f61def072
BLAKE2b-256 checksum
How to use checksums
f02ab805e24db03459a8ecfa32dc164fb381b6043dce9a77aa16502321cf0574
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page