SLM Forge
https://github.com/user-attachments/assets/e9d6d7d1-d623-4441-8c28-ec0e63f40e6e
One-minute tour: describe the model in a sentence. The Tuner picks the smallest model that fits your Mac, finds public data first and asks before every run. It trains and scores the model, and there is an Advanced view with every setting. You export to GGUF or Hugging Face, then chat with the result.
Build your own small language model on a Mac, by talking to an agent.
Describe what the model should do in one sentence. The Tuner, a GPT-6 Luna agent built on the OpenAI Agents SDK, takes it from there. It aims for a small, finished model you enjoy talking to: the smallest base model that can do the job (preferring ones already on your Mac), training data found on Hugging Face (with at most a small set written by GPT-6 to fill gaps), a quick fine-tune with MLX, an honest before/after test, an optional refinement round where GPT-6 reviews and corrects the answers, and an export. It does the groundwork on its own and explains each decision in plain language, but every run waits for your go-ahead: downloading the base model, each training run and the export. At the end you chat with the finished model on the Try it page. You can steer the Tuner at any time by typing, or pause it.
You ──chat──▶ Tuner (GPT-6 Luna) ──tools──▶ models · data · training · evaluation · export
│ (all local, on Apple Silicon)
└──▶ live canvas: every stage rendered as it happens
Training and inference run locally on Apple Silicon. Only the agents call out to an LLM.
The Studio
The main screen is split in two:
- Left: the Tuner. A conversation that streams token by token. The agent's actions appear
as compact steps ("26 steps · searched datasets ×17…"). When the next step is a run, the Tuner
proposes it and a card appears above the input with Go ahead and Not now (a plain
"yes" in the chat works too). Nothing starts until you confirm. Between runs, autopilot keeps
it working on the groundwork, and a finished job wakes it to interpret the result. Autopilot
stops when the Tuner calls
finish_project, or pauses itself (and says so) after two nudges without progress. There's an on/off switch in the top bar. Only creating a project (or pressing Start the Tuner) starts it; opening a project from the sidebar never does. Pausing or stopping a project also cancels the Tuner's current turn and blocks new jobs until you resume it or type a message. - Right: the live canvas. A stage rail (Goal → Data → Model → Train → Evaluate → Refine → Export) above cards that fill in as the work happens: the chosen model with its memory footprint, training sets with cleaning stats, live loss curves, before/after answers, A/B comparisons you judge with one click, and the exported model. The console underneath streams the agent's tool calls and the running job's output.
Once a model is exported, Try it (/p/:id/try, linked from the Studio and Home) opens a chat
with the exported model itself: the fused, quantized folder on disk, exactly as it would run
elsewhere. The page also shows the folder and the command to run it outside SLM Forge.
Finishing doesn't close a project. Keep improving on the Done banner (or just asking the Tuner)
starts another round from the current model, and the new export gets a new name (-v2, or -2
if the name is taken), so earlier models are never overwritten.
Talking to the Tuner never costs anything. Between the export and your next go-ahead, anything that spends (writing examples with GPT-6, an AI review, a download) is a card you confirm, and questions get answers rather than actions.
Every project has the same switch at the top right: Studio, Advanced and ▶ Try it (once a model is exported). For ML experts, Advanced shows the same project with every setting and number: the same stages and ticks as the Studio, each opening a full screen (data, base model, training hyperparameters and runs, evaluation scores and test set, feedback and DPO, export), plus a Playground to chat with any checkpoint.
Quick start
Requirements: an Apple Silicon Mac (M1 or later; 16 GB of memory or more recommended), Python 3.13, and an OpenAI API key for the Tuner.
Install from PyPI (needs uv or pipx):
uv tool install m37labs-slm-forge # or: pipx install m37labs-slm-forge
echo "OPENAI_API_KEY=sk-..." > .env # in the folder you start it from; add HF_TOKEN=hf_... to publish
slm serve # → http://127.0.0.1:8000
The package is m37labs-slm-forge on PyPI (plain slm-forge is too close to an existing project
name). It installs the slm command; uvx m37labs-slm-forge serve runs it without installing.
An installed copy keeps its projects, data and models in
~/Library/Application Support/SLM Forge/ (set SLM_WORKSPACE to move them).
Where your keys go: put them in ~/Library/Application Support/SLM Forge/.env to use them
wherever you start the app. A .env in the folder you run slm serve from also works, and wins
if a key is set in both. Keys exported in your shell override both files.
Or run from source (also needs Node 20+):
git clone https://github.com/engagepy/SLM-Forge.git && cd SLM-Forge
uv sync # Python deps (MLX, mlx-lm, mlx-lm-lora, FastAPI…)
cp .env.example .env # then add OPENAI_API_KEY
npm --prefix web ci && npm --prefix web run build # the web UI
uv run slm serve # → http://127.0.0.1:8000
A clone keeps its data in ./workspace. The Tuner uses the model named in SLM_OPENAI_MODEL
(default gpt-6-luna); set it in .env if your OpenAI account uses a different model.
Useful commands:
uv run slm hardware # what this Mac can train
uv run slm models qwen # base models that fit, with memory estimates
uv run python scripts/smoke.py # end-to-end check: download → SFT → DPO → export (~4 min)
uv run pytest # unit + API tests (no GPU or network needed)
For UI development, run uv run slm serve and npm run dev in web/ (Vite proxies /api).
The loop
| Step | What happens |
|---|---|
| 1. Base model | Search Hugging Face (MLX conversions first). Each model gets a fit verdict and a memory breakdown for inference and each training preset, measured against this Mac's GPU budget. |
| 2. Data | DataScout searches the Hub, reads dataset cards, previews rows and proposes imports with a column mapping. When data needs a login or lives elsewhere, it writes step-by-step instructions and gives you an upload slot. You can also browse or upload yourself. DataPrep proposes the mapping and cleaning rules. Preparing a dataset normalises, dedupes, filters, splits and measures token lengths into an immutable version. |
| 3. Train (SFT) | LoRA / DoRA / full fine-tuning with every mlx-lm knob exposed. Presets are sized to your hardware using the data's real token lengths, with a live memory estimate before launch and live loss/LR curves during the run. Epochs are converted to iterations for you. |
| 4. Playground | Chat with any checkpoint or the base model. Every sampling control is available: temperature, top-p/k, min-p, repetition/presence/frequency penalties, XTC, seed, thinking mode. |
| 5. Feedback | Compare two answers, pick one, optionally rewrite it, add a critique. Picks become DPO preference pairs; rewrites become SFT examples too. Observer reads the feedback and training results and proposes the next step. Synth writes new examples aimed at the weaknesses your critiques describe. Its preference pairs use your model's own answer as the rejected side when the GPU is free. |
| 6. DPO | Preference-tuning on your feedback via mlx-lm-lora. The current SFT model is fused first so it serves as the frozen reference. |
| 7. Export | Fuse, optionally quantize, and write a model card with the model's real lineage and the minimum Mac memory it needs. Runs anywhere with mlx_lm.generate --model <path>. |
Agents only propose. Anything that downloads, trains or adds data waits in the Agents page for your approval (and you can edit the payload first). Synthetic examples wait in a review queue before they can reach training.
"RLHF" here means DPO, not PPO. PPO needs the policy, a reference and a reward model in memory at once, which doesn't fit a 16 GB Mac for useful model sizes. DPO learns from the same human preference pairs with just the policy and a frozen reference.
Agent providers
The Tuner and its specialists (DataScout, DataPrep) run on the
OpenAI Agents SDK with streaming and persistent
session memory. SLM_AGENT_PROVIDER picks the model behind them, and behind the AI judge and the
Advanced screens' agents too:
| Provider | Key | Model | How it runs |
|---|---|---|---|
openai (default) |
OPENAI_API_KEY |
gpt-6-luna (SLM_OPENAI_MODEL) |
Responses API. |
claude |
ANTHROPIC_API_KEY |
claude-fable-5-1 (SLM_CLAUDE_MODEL) |
The Agents SDK's LiteLLM extension for the Tuner; the Anthropic SDK for the judge. |
ollama (experimental) |
none | qwen2.5:7b-instruct (SLM_OLLAMA_MODEL) |
Ollama's OpenAI-compatible endpoint. Fully offline, but a small local model rarely drives the Tuner's long, tool-heavy loop well, and it competes with training for memory. |
The prompts and tools are the same for every provider; only the model changes. Agent traces go to
the OpenAI dashboard only with the OpenAI provider and only if you opt in with
SLM_OPENAI_TRACING=true. The spend meter reads OpenAI's billing only.
Lessons the Tuner carries
Its instructions encode what real runs on a 16 GB Mac taught, and the code enforces the same
things: scored test sets sized by task (exact match for deterministic outputs), one output format
per project, exports that carry their system prompt, no tiny top-ups, no DPO under 30 pairs,
learning rate 1e-4, continued runs that resume the adapter instead of copying the model, and
spending only inside a round the user set in motion. AGENTS.md lists each with the run that
taught it.
How the Tuner works
- One agent, 33 tools (
src/slm/tuner/tools/) wrapping the tested platform code:get_status,find_base_models,choose_base_model,search_datasets,prepare_dataset,generate_synthetic_examples,start_training,try_model,ask_user_to_compare,export_modeland more. Quick jobs (imports, data prep, synthesis) are awaited inside the tool. Runs (choose_base_model,start_training,export_model) only record a proposal (tuner/confirm.py); your confirmation submits the job, and the finished job wakes the Tuner with a[Job update]turn. - Its instructions carry what running this platform taught us: learning rate 1e-4 is safe
and 2e-4 diverged; near-duplicates make validation loss meaningless;
max_seq_lengthmust cover the data's p95 tokens; when public data is poor, write a small seed set, never the dataset. - Public data first, small synthetic sets only: the Tuner's DataScout hunts the Hugging Face Hub (and the alternatives it can reach), imports generously and samples down to the plan's target. The teacher model writes at most 50 examples per call and 200 per project: a seed of a few dozen when nothing public fits, or a top-up aimed at a gap the evaluation showed. It never writes the dataset itself.
- Before/after is a number. The Tuner writes a test set sized to the task (10–80 cases, with
exact expected outputs wherever the task is deterministic) and scores every checkpoint on it with
evaluate_model: exact match first, the judge (0–10 against the goal) only where there is no expected output or the answer differs from it. The scores sit on the Evaluate card, the export is the best-scoring checkpoint (serve_checkpointrolls back if the latest run made things worse), and DPO isn't attempted under 30 pairs. - Try it keeps you testing: four suggested inputs sit above the message box, three on-goal at varied difficulty and one that should get the model's empty or negative answer (dashed). Used ones stay ticked; once all four are used, GPT-6 writes four fresh ones scoped to the goal.
- Exports work with no flags. The project's system prompt is built into the exported model's
chat template, so
mlx_lm.generate --model <folder> --prompt "…"(or any loader) behaves like the app. Exports made before this show a--system-promptin their command instead. - A metrics strip above every project: the Mac and its ML budget, the GPU, how much disk the app occupies (datasets, runs, exports, uploads, databases and the downloaded models, with free space), and the OpenAI spend below.
- Disk is not the constraint. The Tuner keeps every adapter, metric, evaluation and example,
and sizes data by the goal (hundreds for a persona, thousands for extraction or JSON). What runs
leave behind that is reclaimable, the fused model copies each run writes for the next and the
folders of failed runs, shows in the Disk pill; past
SLM_DISK_TIDY_GB(20) it's suggested, and one click clears it with rollback intact (checkpoints re-fuse from their adapters). - Storage & cleanup (
/storage, from the Disk pill or the sidebar): every base model the app downloaded (and which projects use it), every project with the size of its runs, data and exported models. Remove a model, delete an export, or delete a project (keeping its exports if you like); each shows the space it freed and the meter updates at once. - A spend meter, read from OpenAI. The pill in the Studio bar and the sidebar shows what the
OpenAI account has spent today and this month, straight from OpenAI's Costs API (no token
counting). It needs an organisation admin key in
.env(OPENAI_ADMIN_KEY; the project key can't read costs). It finds the OpenAI project your key belongs to by itself (OPENAI_PROJECT_IDoverrides). OpenAI updates the figure with a lag of a few hours; the meter refreshes every ten minutes. OpenAI only, for now. An admin key can read and manage your whole organisation, so it is optional: create a dedicated one, or leave it unset and the meter reads "spend not set up". - AI feedback instead of human clicks:
ai_review_answershas the local model answer each prompt twice, and GPT-6 picks the better answer, writes the ideal one and critiques the flaws. Each verdict becomes a DPO preference pair and, where the model was wrong, a corrected SFT example. Human A/B judging is still available when you ask for it. - Base models built to be small.
find_base_modelslists a curated catalog first (Qwen2.5, Qwen3, SmolLM2/3, Llama 3.2, Gemma 3, Granite, Phi-4 mini) with licence, what each is best for and why, ordered by the plan's task type; the Tuner shortlists 2–3 and proposes one. - Specialists for data. The Tuner delegates the hunt to DataScout (an Agents-SDK agent that searches from several angles, previews candidates in parallel and returns a ranked shortlist with licence, mapping and fit score) and the cleaning plan to DataPrep (inspects rows, checks a mapping against them, returns thresholds and sequence length). Their tool calls show in the console under their names; the Tuner's own context stays small.
- Big imports, then a sample. Hub datasets stream in up to 200,000 rows with progress on the
canvas;
prepare_dataset(max_examples=…)cleans, deduplicates and samples to the plan's target, and the rest stays on disk as a pool for later rounds. - Smallest model that does the job: about 0.5B by default, 1–1.5B when answers need real explanation, 3B only if you ask or a smaller model has clearly failed. Models already in the local Hugging Face cache are listed first and registered instantly, with no re-download.
- Turns run on one long-lived event loop. The SDK's shared OpenAI client binds to the first
loop it runs on, so a fresh
asyncio.run()per turn fails with "Event loop is closed". Tools run in worker threads so blocking work never stalls the stream. Messages that arrive mid-turn (from you, a job, or finished comparisons) batch into the next turn. - Memory: the SDK's
SQLiteSessionkeeps the agent's context across turns and server restarts; the visible transcript is stored separately.
Architecture
web/ (React + Vite + TanStack Query + Recharts) ──REST + SSE──▶ FastAPI src/slm/api
│
models/ hub search, fit check, download agents/ provider (OpenAI | Claude | Ollama)
data/ scout tools, mapping, cleaning, splits scout · prep · observer · synth
train/ configs & presets, subprocess runner, actions: approved proposal → job
log parsing, diagnosis, 3-lane job worker
inference/ single resident model, streaming generation export/ fuse, quantize, model card
SQLite (SQLModel) + workspace/ on disk
Design decisions worth knowing:
- Training runs as a subprocess (
mlx_lm lora,mlx_lm_lora.train) from a generated YAML config, and its log is parsed into metrics. A crash or OOM can't take down the server, and all GPU memory comes back when the job ends. - Three job lanes, one job each:
gpu(training and export; evicts the chat model first),io(downloads, imports, data prep) andagent(LLM agent runs). - All in-process MLX work runs on one dedicated thread. MLX's thread-local compile cache holds Python objects; on the main thread it's destroyed after the interpreter shuts down and segfaults the process on exit.
- Runs build on each other without copying the model. A new SFT run continues the served adapter on the shared base model (no fused copy per run; a project's history is small adapters), unless you start from the base model to compare settings fairly. Only a DPO round and an export write a fused model. Every checkpoint records its parent; exports list only the served model's ancestry.
- Every finished run is diagnosed for divergence, overfitting, no improvement and "memorised" validation (near-zero validation loss usually means the validation set overlaps training). Warnings show on the run page and feed the Observer.
What's been verified (M1 Pro, 16 GB)
With mlx-community/Qwen2.5-0.5B-Instruct-4bit:
- Full loop through the API: SFT → A/B feedback → DPO → export. The exported model runs standalone at ~245 tok/s in 0.35 GB.
- Memory estimate: 1.37 GB estimated vs 1.48 GB measured for a balanced-preset SFT run.
- Learning rate: 2e-4 diverges (loss spikes 1.9 → 7.9) and is flagged automatically; 1e-4 trains cleanly.
Roadmap
Built since the first plan: the scored evaluation harness (exact match + judge), checkpoint comparison and roll-back, adapter resumption between rounds, the storage manager, the spend meter, the model catalog and the data specialists. Still ahead:
- near-duplicate detection (MinHash), language ID and PII scrubbing in cleaning;
- web search for off-Hub data;
- a machine-wide GPU lock across server processes, and a memory estimator recalibrated for 3B+;
- curricula, ORPO/GRPO, GGUF/Ollama export, a device compatibility matrix.
Share it: GGUF and Hugging Face
Every export in the Studio has two more buttons:
- Make GGUF converts the model (exactly as exported, built-in system prompt included) into GGUF
files for llama.cpp, Ollama and LM Studio: Q4_K_M (small, the usual choice) and Q8_0
(near-lossless). The first time, SLM Forge downloads llama.cpp's converter and its dependencies
(about 300 MB, into the workspace). Q4_K_M needs llama.cpp's quantizer:
brew install llama.cpp. Without it you get Q8_0. - Upload to Hugging Face publishes the export (the MLX model, the GGUF files, the model card and
the base model's licence files) to a repository under your account, public unless you choose
private. It uses your Hugging Face login:
HF_TOKENin.env(a token that can write) oruv run hf auth login. The form shows the base model's licence first; Llama models are only published with Meta's licence file, and nothing that names a folder on your Mac is uploaded. Afterwards:ollama run hf.co/<you>/<model>:Q4_K_M.
The Tuner can do both too, as cards you confirm, when you ask it to share the model.
Privacy & data
Training and inference run on your Mac. Three things leave it:
- OpenAI (with the default provider: the Tuner, its specialists and the AI judge): your goal and chat, samples of your
dataset rows, the test set, and the local model's answers when they are scored or reviewed. Don't
put data you may not share with OpenAI through it, such as health records or other personal
data. Agent traces go to your OpenAI dashboard only if you set
SLM_OPENAI_TRACING=true. - Hugging Face: model and dataset downloads and dataset searches, and the models you choose to
publish. It uses
HF_TOKENfrom.envor yourhf auth login; SLM Forge never stores it elsewhere. - GitHub and PyPI, once, when the first GGUF export downloads llama.cpp's converter.
- Anthropic instead of OpenAI, the same data, if you set
SLM_AGENT_PROVIDER=claude. Withollamanothing leaves the Mac for the agents.
API keys stay in .env or your environment and are never logged. The server listens on
127.0.0.1 only and has no login: don't expose it to a network. There is no telemetry.
Licences
SLM Forge is released under the Apache License 2.0; third-party material in this repository is listed in NOTICE.
The models you build are yours to use, within the terms of what they are made from:
- The base model's licence. Each candidate shows its licence and conditions before you confirm it. Qwen2.5-3B is research-only; Llama 3.2 needs "Built with Llama" and follows Meta's acceptable use policy; Gemma passes its use restrictions on. Every export copies the base model's licence files and writes the licence, its conditions and the required notices into its model card.
- The training data's licences. Imported datasets keep their licence, and the model card lists each one. Check that a dataset's licence allows your use before you train on it.
- OpenAI's terms, for examples the teacher model wrote or corrected. The model card says when any were used.
This is a summary to help you check, not legal advice.
Contributing
Contributions are welcome: read CONTRIBUTING.md (setup, checks, and the invariants in AGENTS.md) and the Code of Conduct. Report security problems privately, as described in SECURITY.md.
Metadata
Release files for m37labs-slm-forge 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| m37labs_slm_forge-0.2.0.tar.gz | 424.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| m37labs_slm_forge-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 888.1 kB
Release files / m37labs_slm_forge-0.2.0.tar.gz
| Download URL | m37labs_slm_forge-0.2.0.tar.gz |
|---|---|
| Size | 424.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
982615432e72a6e65c9278a1978165f5ca9de2e435f161dda58c5ceda2dffc5a
|
|
BLAKE2b-256 checksum How to use checksums |
b43ccaa8eac4720b9a11a739f5dcfd33a5e985dfa6332fc3d137dc6a1d9cf082
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / m37labs_slm_forge-0.2.0-py3-none-any.whl
| Download URL | m37labs_slm_forge-0.2.0-py3-none-any.whl |
|---|---|
| Size | 464.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
690d0aa7efd19b36fe531217cc11b98b84cef100d1e6a2a44f1900d7d78a3fe7
|
|
BLAKE2b-256 checksum How to use checksums |
2bbf19d7610ad2493786614076e0c72b9702a2ee67e42d6bfde8abbe9a8da20b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log