tinyforge
Train, evaluate and serve small LLMs from scratch on modest GPUs (developed against a 4 GB RTX 3050 Ti). One engine, three front doors: CLI (for developers/CI), REST API, and a web UI (which can later be wrapped as a desktop app with Tauri/Electron without changing the backend).
Status: early alpha (0.1). Tested on Windows and on Linux (a Colab T4 GPU), each with one NVIDIA GPU; multi-GPU and
Kubernetes-with-GPU are not tested yet. Step-by-step Linux, Windows and Colab instructions: docs/install.md.
Install: pip install tinyforge, then tinyforge doctor to see what else you need
(PyTorch, pip install "tinyforge[finetune]", and optional pieces such as Soup and llama.cpp).
Quick start
The same commands work on Linux and Windows (venv activation, environment variables and curl differ: see
docs/install.md for both).
python3 -m venv .venv && source .venv/bin/activate # Windows: py -3 -m venv .venv ; .venv\Scripts\Activate.ps1
pip install torch --index-url https://download.pytorch.org/whl/cu124 # or /cpu
pip install tinyforge # or, from a clone: pip install -e ".[dev]"
tinyforge doctor # hardware + warnings, with fixes
tinyforge pipeline --preset micro --steps 2000 # doctor -> data -> plan -> train -> eval
tinyforge generate "ROMEO:" --int8
tinyforge serve # UI at http://127.0.0.1:8000
What you get
| Stage | What it does | Guard rails (warning codes) |
|---|---|---|
doctor |
GPU/VRAM/bf16/Triton/JAX probe | HW001-HW008 |
data prepare |
download or ingest text, train BPE, contiguous train/val split | DQ001-DQ005: tiny corpus, duplicates, bad vocab |
plan |
VRAM estimate; auto-enables grad checkpointing, shrinks micro-batch, keeps effective batch | PL001-PL005 |
train |
AMP (bf16/fp16+scaler), grad accumulation, cosine LR, atomic checkpoints, auto-resume | TR001 NaN skip/abort, TR002 overfit, TR003 plateau |
eval |
perplexity vs random baseline, diversity, memorisation, determinism, KV-cache correctness, speed/VRAM | EV001-EV009, non-zero exit on failure |
generate/serve |
KV-cache decoding, top-k/top-p, int8 weights, optional Triton RMSNorm |
Fine-tuning platform commands
| Command | What it does |
|---|---|
memory |
probe VRAM/RAM/disk and say where a model's frozen weights would live and the estimated speed |
data ingest |
documents (txt, md, html; pdf/docx/pptx/xlsx with the docs extra) -> cleaned, chunked, de-duplicated JSONL |
data tabular |
CSV -> text-to-SQL examples whose answers are verified by executing them (--hard for eval shapes) |
data pairs |
chunks -> grounded Q&A pairs from a teacher model (any OpenAI-compatible endpoint), split before generation |
data pii-scan, --pii |
find emails, phones, cards, IDs and credentials; flag, redact or drop |
worker run --spec job.yaml |
run a TrainingJob; backend: native or soup (layer streaming for models larger than VRAM) |
export gguf |
adapter + base -> GGUF for llama.cpp |
bench list/suggest/run |
pick benchmarks sized to your GPU/RAM/time budget (MMLU, ARC, HellaSwag, GSM8K, IFEval via lm-evaluation-harness, plus a built-in SQL execution check) and run them, base vs tuned |
serve [--engine llamacpp] |
API + UI; OpenAI-compatible /v1/chat/completions with guardrails and metrics |
Verified on one RTX 3050 Ti 4 GB laptop (Windows): a 3B and an 8B model fine-tuned through the job spec and served
locally; see benchmarks/ and docs/soup-backend.md. Early alpha: single GPU; Windows and a Linux Colab T4 tested.
Model: decoder-only transformer with RoPE, RMSNorm, SwiGLU, tied embeddings, PyTorch SDPA (FlashAttention kernels).
Presets: nano ~1M, micro ~12M, small ~28M, base ~100M.
Fine-tune a pretrained model (LoRA / QLoRA)
The most useful thing a small GPU can do. Default: SmolLM2-360M-Instruct on an instruction dataset.
pip install -e ".[finetune]"
tinyforge ft pipeline --steps 150 # doctor -> data -> plan -> train -> eval + merge
tinyforge ft generate "Explain hash tables" # tuned model; add --base to compare with the original
| Stage | What it does | Guard rails |
|---|---|---|
ft data |
JSONL or HF dataset -> chat JSONL; dedupe; hash split keyed on prompt (no leakage); length stats | FD001-FD006: too few examples, duplicates, malformed, tiny val set, truncation, PII/credentials (credentials are dropped) |
ft plan |
picks fp16 vs 4-bit NF4, checkpointing, and a token budget per micro-batch to fit VRAM | FP001-FP007 |
ft train |
LoRA, length-grouped token-budget batching, resume, best-adapter tracking | FT001 NaN, FT002 overfit, FT005 VRAM spill, FT006 adaptive OOM recovery |
ft eval |
tuned vs base held-out loss, forgetting check on general text, side-by-side generations, merge + equivalence check | FE001-FE009; exit 1 if tuned is not better than base |
Measured on an RTX 3050 Ti Laptop (4 GB), SmolLM2-360M, 150 steps x 16 examples: 2.4 GB peak VRAM, ~830 tok/s,
~10 min, held-out loss 1.314 -> 1.262 (-4.0%), no forgetting; exports a 690 MB standalone merged model.
Lessons baked into the planner (all measured, see finetune.plan): small micro-batches are launch-bound (bs8 is
2.3x the tokens/s of bs4), LoRA dropout 0 is ~40% faster, and allocator-level estimates miss cuBLAS workspace, so
training recovers from memory errors by shrinking the token budget instead of crashing.
Layout
src/tinyforge/ engine (config, data, model, train, evaluate, infer, kernels/), cli.py, server.py, ui/
tests/ unit + end-to-end smoke (synthetic data, no network)
infra/terraform/ S3 artifacts, IAM, no-inbound GPU spot worker (SSM access), budget alarm
.github/ CI: lint, test matrix, terraform validate, docker build, self-hosted GPU e2e
Dockerfile CUDA runtime image
Cloud (AWS)
cd infra/terraform
terraform init -backend-config="bucket=<state-bucket>" -backend-config="key=tinyforge/tf.tfstate"
terraform apply -var enable_gpu_worker=true -var alert_email=you@example.com
# then run the `connect` output command to port-forward the UI over SSM (no open ports)
Verified results (what was actually run)
Unless a row says otherwise it was run on one machine: Windows 11, RTX 3050 Ti Laptop GPU (4 GB), 16 GB RAM. Each row links to the raw numbers. These are single runs on small test sets, not benchmarks: read the caveats in each JSON file.
| What was tested | Result | Data |
|---|---|---|
| Llama-3.1-8B fine-tuned on a 4 GB GPU (layer streaming through the job spec), served with llama.cpp | trained 200 steps in 9.5 min; on 100 questions about a table it never saw: 100% execution accuracy vs 93% for the same model told to reply with SQL only | e8-penguins-8b-soup.json |
| Same 8B model, harder question shapes never used in training | 98% vs 97%: no meaningful gain over a well-prompted base model. Fine-tuning mostly fixed output format | e8-penguins-8b-hard.json |
| Qwen2.5-3B fine-tuned end to end (data -> train -> evaluate -> serve) | 98% execution accuracy vs 70% for the base model told to reply with SQL (0% without that instruction) | e2e-titanic-3b-soup.json |
| Soup-trained adapter, quality check (3B) | exact match 4% -> 45%, SQL that runs 9% -> 88% on 100 held-out rows | soup-adapter-check-qwen3b.json |
| Chunked cross-entropy (memory saving), Qwen2.5-1.5B 4-bit | peak GPU memory 3.65 -> 2.95 GB, 289 -> 499 tok/s, same loss | chunked-ce-qwen1p5b-4bit-rerun.json |
| llama.cpp inference options, 3B | flash attention ~8% faster; 8-bit KV cache halves the cache (144 -> 76.5 MB); n-gram speculation 129 vs 54.5 tok/s only on repeated requests (no gain on a first request) | inference-techniques-qwen3b.json |
| A deliberately bad model (random weights, real architecture) | the tool now fails it: FE010 base model looks untrained, FE011 degenerate looping output (before the fix it printed "passed"). A model without a chat template gets a clear FD007 error |
tests/test_bad_models.py |
| Test suite | 294 tests pass on Windows (Python 3.12) and on Linux in Docker/CI (Python 3.10 and 3.12, CPU PyTorch) | pytest -q |
| Linux + real GPU (Colab, Tesla T4 16 GB, Python 3.13), installed from PyPI | doctor, data tools, from-scratch pipeline (all quality gates pass), LoRA fine-tune (held-out loss 1.314 -> 1.271), job-spec worker, server with OpenAI API and secret guardrail, GGUF export and llama.cpp serving all worked. Three bugs found there (benchmark runner on newer transformers, a missing-checkpoint crash, torchao message) are fixed on main and ship in the next release; the in-notebook test-suite run still needs a clean re-run |
notebooks/colab_smoke_test.ipynb |
What these results do not show: general-purpose quality gains (the SQL tests use templated questions), results on
other hardware or operating systems, multi-GPU training, or the standard benchmark runner (bench run, built on
lm-evaluation-harness), which has not been run end to end yet.
The dashboard while serving a 3B model. KV-cache use grows linearly with context (36 KiB per token) while GPU memory stays flat, because llama.cpp reserves the cache up front.
Honest limits / roadmap
- JAX: there is a probe and a backend slot, but training is PyTorch-only today. JAX has no native Windows GPU
support (use WSL2/Linux), and a second backend should be added behind the same
train()interface. - Triton: one fused kernel (RMSNorm forward, inference only). Next: fused SwiGLU, then a backward pass.
Triton on Windows needs the
triton-windowspackage. - From-scratch models are for learning/prototyping; use fine-tuning for anything useful. A 4% held-out gain on a general instruction set is modest by design: LoRA on a small, already-tuned model mostly shifts style. Use a task-specific dataset to see larger gains.
- 4-bit QLoRA is verified on real hardware (see ROADMAP.md): it works but ends ~4.7% worse in loss than fp16 on a
model that fits in fp16. Its value is fitting larger models: an 8B model trains on a 4 GB GPU via
backend: soup. - Generation speed (HF
generate, eager) is ~10 tok/s on this GPU; the merged model is the thing to export to llama.cpp/vLLM for serving. - Single-job (a SQLite job queue exists). Multi-GPU: data-parallel LoRA on one machine (
tinyforge ft train --gpus N, orresources.gpus: Nin a job spec) is implemented and tested with real 2-process training on CPU (gloo) on Windows and Linux, and on real GPUs with NCCL (Kaggle 2x T4,notebooks/kaggle_multi_gpu.ipynb): both GPUs stay identical and resume works, but throughput is only about 1.2x-1.4x of one GPU for a small 360M LoRA job, well short of 2x (cause being investigated). Not built: FSDP/sharded training (models too big for one GPU), multi-node, pipeline parallelism. - The API requires
TINYFORGE_API_TOKEN(bearer) and the server binds to localhost by default; the static UI page itself is not authenticated. Put it behind TLS, SSM/VPN or a reverse proxy before exposing it.
Metadata
Release files for tinyforge 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tinyforge-0.1.1.tar.gz | 191.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tinyforge-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 345.9 kB
Release files / tinyforge-0.1.1.tar.gz
| Download URL | tinyforge-0.1.1.tar.gz |
|---|---|
| Size | 191.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4cc3a67eed10290d2d9fc86da5e433fed67c9459346f5df5f958e851414cd34c
|
|
BLAKE2b-256 checksum How to use checksums |
396442dfd0e8a9af7dbbbf08f0a27deebe5315c5854cc21f05b70b7e59b40a0f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / tinyforge-0.1.1-py3-none-any.whl
| Download URL | tinyforge-0.1.1-py3-none-any.whl |
|---|---|
| Size | 154.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bd4038c8d0ee0e8069b063976209fc06d0fb490ee120b23d7989e80738025005
|
|
BLAKE2b-256 checksum How to use checksums |
a605c7f3c40d969f6f81129e5b76d7860dd2381a2a64c6f6df53e26e8b5d5c47
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log