SlickTune 🧩: Composable LLM Fine-Tuning by SlickML🧞
Explore Releases 🟣 Become a Contributor 🟣 PyPI 🟣 Join our Slack 🟣 Tweet Us
🧠 Philosophy
SlickTune 🧩 is a small, composable toolkit for teaching LLMs new facts and behaviors with Transformers + PEFT + TRL. LoRA / QLoRA are PEFT adapters; full FT updates every weight. The goal is the same SlickML spirit: prototype fast 🏎, keep axes orthogonal, and measure whether the model actually learned your facts 🔎.
New to fine-tuning? Start here → Fine-Tuning LLMs: A Visual Guide — pre-training vs prompting vs FT, Full / LoRA / DoRA / AdaLoRA / QLoRA with diagrams, multi-adapter merge (TIES / DARE), how to choose a strategy, and probes & holdout PPL.
Fine-tuning is an orthogonal stack — swap any axis without rewriting the others:
model × strategy × objective × data × metrics
🧩 Abstractions
flowchart TB
subgraph inputs [Inputs]
modelId[model_id]
dataJsonl[data JSONL]
end
subgraph axes [Composable axes]
strategyNode["Strategy: LoRA / DoRA / AdaLoRA / QLoRA / Full"]
objectiveNode["Objective: SFT / DPO / ORPO / KTO / GRPO"]
end
subgraph core [Tuner fit]
tuner[Tuner]
loadStep[load model and tokenizer]
applyStep[strategy.apply]
trainStep[TRL trainer]
metricsStep[MetricsTracker]
end
subgraph outputs [Outputs]
checkpoint[adapter or checkpoint]
metricsFile[metrics.json]
probeRate[probe pass rate]
end
modelId --> tuner
dataJsonl --> tuner
strategyNode --> tuner
objectiveNode --> tuner
tuner --> loadStep --> applyStep --> trainStep --> metricsStep
trainStep --> checkpoint
metricsStep --> metricsFile
checkpoint --> probeRate
| Axis | Responsibility | Shipped (phases 0–5) |
|---|---|---|
| Strategy | How weights change (PEFT vs full) | LoRA / DoRA / AdaLoRA / QLoRA / Full |
| Objective | What is optimized / data contract | SFT / DPO / ORPO / KTO / GRPO |
| Data | Examples → chat, prefs, or rewards | train + holdout + prefs/KTO/GRPO JSONL (about_amir*.jsonl) |
| Metrics | Comparable run stats | MetricsTracker (+ holdout PPL, judge score) |
| Merge | Combine / bake adapters | TIES / DARE / linear + bake_adapter |
| Eval | Holdout + judges | slicktune eval, SubstringJudge, LLMJudge |
| Probe | Did the model learn your facts? | slicktune probe |
📌 Quick Start
from slicktune import LoRAStrategy, SFTObjective, Tuner
Tuner(
model_id="HuggingFaceTB/SmolLM2-135M-Instruct",
strategy=LoRAStrategy(r=8),
objective=SFTObjective(),
output_dir="outputs/sft_lora",
eval_data="examples/data/about_amir.eval.jsonl",
).fit("examples/data/about_amir.jsonl")
👤 Personal “about me” loop (recommended)
- Edit
examples/data/about_amir.jsonlwith facts about you (or keep the SlickML starter facts) ✍️. - Edit
examples/data/about_amir.eval.jsonlwith held-out paraphrases (same topics, not copied from train) for perplexity 📉. - Edit
examples/data/about_amir.probes.jsonlwith questions and amust_containsubstring that should appear after training 🎯. - Train a strategy on a tiny instruct model 🧪.
- Probe the checkpoint — pass rate shows whether fine-tuning stuck ✅.
before FT → model guesses / hallucinates about you
after FT → probe answers contain your facts
🛠 Installation
Install Python >=3.10,<3.14 and uv, then simply run 🏃♀️:
uv sync --locked --all-extras --all-groups
QLoRA (CUDA + bitsandbytes only) 🔥:
uv sync --extra qlora
Task runner is Poe the Poet (same idea as slick-ml, with uv instead of Poetry). Install the CLI once 🏃♀️:
uv tool install poethepoet
poe greet
Developer workflow (format / check / test) lives in CONTRIBUTING.md 🧑💻🤝.
🚂 Train each strategy
Default demo model: HuggingFaceTB/SmolLM2-135M-Instruct (small enough for laptop smoke tests) 💻.
🟢 LoRA + SFT (default — works on Mac MPS / CPU / CUDA)
uv run slicktune train \
--strategy lora \
--data examples/data/about_amir.jsonl \
--eval-data examples/data/about_amir.eval.jsonl \
--output outputs/sft_lora \
--epochs 20
uv run slicktune probe \
--model-dir outputs/sft_lora \
--probes examples/data/about_amir.probes.jsonl
Or: poe train-lora / poe probe-lora / poe eval-lora / uv run python examples/run_sft_lora.py
🟣 DoRA + SFT
uv run slicktune train \
--strategy dora \
--data examples/data/about_amir.jsonl \
--eval-data examples/data/about_amir.eval.jsonl \
--output outputs/sft_dora
# or: uv run python examples/run_sft_dora.py
🟤 AdaLoRA + SFT
uv run slicktune train \
--strategy adalora \
--data examples/data/about_amir.jsonl \
--eval-data examples/data/about_amir.eval.jsonl \
--output outputs/sft_adalora
# or: uv run python examples/run_sft_adalora.py
🟡 LoRA + DPO (preference pairs)
uv run slicktune train \
--strategy lora \
--objective dpo \
--data examples/data/about_amir.prefs.jsonl \
--output outputs/dpo_lora \
--epochs 10
# or: poe train-dpo / uv run python examples/run_dpo_lora.py
🟢 LoRA + KTO (unpaired labels)
uv run slicktune train \
--strategy lora \
--objective kto \
--data examples/data/about_amir.kto.jsonl \
--output outputs/kto_lora \
--epochs 10
# or: poe train-kto / uv run python examples/run_kto_lora.py
ORPO: --objective orpo with the same prefs JSONL as DPO (TRL experimental).
🔴 LoRA + GRPO (verifiable substring rewards)
GRPO samples multiple completions per prompt and scores them with a verifiable
must_contain reward (exact match = 1.0, else keyword-overlap fraction). On a
cold tiny base model rewards stay ~0 so GRPO cannot learn — warm-start with SFT
first. The CLI train command has no --adapter-path; use the smoke example
(or Tuner(..., adapter_path=...) in Python):
# SFT warm-start → GRPO (writes outputs/grpo_lora_sft then outputs/grpo_lora)
poe train-grpo
# or: uv run python examples/run_grpo_lora.py
🟣 Merge adapters (TIES / DARE)
Combine multiple PEFT adapters on the same base (TIES, DARE, linear, …), or bake into full weights for serving engines that want a single checkpoint. Smoke demo trains two tiny adapters then merges them:
poe merge-ties
# or: uv run python examples/run_merge_ties.py
# → outputs/merge_a_lora + outputs/merge_b_lora → outputs/merged_ties
uv run slicktune merge \
--model HuggingFaceTB/SmolLM2-135M-Instruct \
--adapter outputs/merge_a_lora \
--adapter outputs/merge_b_lora:0.5 \
--method ties \
--density 0.5 \
--output outputs/merged_ties
# bake into full weights: add --bake
# alternate: merge any two trained adapters, e.g. outputs/sft_lora + outputs/dpo_lora:0.5
from slicktune import AdapterRef, merge_adapters
merge_adapters(
model_id="HuggingFaceTB/SmolLM2-135M-Instruct",
adapters=[
AdapterRef(path="outputs/merge_a_lora", name="a", weight=1.0),
AdapterRef(path="outputs/merge_b_lora", name="b", weight=0.5),
],
output_dir="outputs/merged_ties",
method="ties",
density=0.5,
)
🔎 Eval harness (holdout PPL + judges)
uv run slicktune eval \
--model-dir outputs/sft_lora \
--eval-data examples/data/about_amir.eval.jsonl \
--probes examples/data/about_amir.probes.jsonl \
--judge substring
--eval-data should be a holdout SFT JSONL (not the training file). The shipped
about_amir.eval.jsonl paraphrases the same topics for holdout perplexity.
Use --judge llm to score generations with an LLM rubric (0–10 → normalized).
On the tiny demo model, prefer --judge substring: the same 135M checkpoint is a
weak judge and will under-score even when answers are correct.
🔵 QLoRA + SFT (CUDA required)
uv sync --extra qlora
uv run python examples/run_sft_qlora.py
On Apple Silicon, use LoRA instead — bitsandbytes 4-bit needs CUDA 🍎.
🟠 Full fine-tuning + SFT
uv run python examples/run_sft_full.py
Heavier on memory; prefer LoRA for iteration 💾.
📦 Data formats
SFT JSONL (any of these per line) 📝:
{"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}
{"prompt":"...","response":"..."}
{"instruction":"...","input":"...","output":"..."}
Probe JSONL 🕵️:
{"prompt":"Who is Amirhessam Tahmassebi?","must_contain":"SlickML"}
Holdout eval JSONL (same SFT shapes as train; keep examples out of the train file) 📉:
{"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}
Ship example: examples/data/about_amir.eval.jsonl.
Preference JSONL (DPO / ORPO) ⚖️:
{"prompt":"...","chosen":"...","rejected":"..."}
KTO JSONL (unpaired labels) ✅❌:
{"prompt":"...","completion":"...","label":true}
GRPO JSONL (verifiable RL; solution is accepted as an alias for must_contain) 🎯:
{"prompt":"Who is Amirhessam Tahmassebi?","must_contain":"founder of SlickML"}
🗺 Roadmap
| Phase | Scope |
|---|---|
| 0–1 | Skeleton, SFT + LoRA/QLoRA/full, metrics, personal probe loop |
| 2 (done) | DoRA / AdaLoRA, holdout PPL + substring/LLM judges |
| 3 (done) | DPO / ORPO / KTO |
| 4 (done) | GRPO / verifiable RL |
| 5 (done) | Merge (TIES/DARE), multi-adapter |
| 6 (now) | Classic RLHF: reward model + PPO (SFT → RM → online PPO) |
| 7 | Optional multimodal PEFT (VLM SFT + LoRA) |
Planned (not implemented yet): Phase 6 adds RewardObjective + PPOObjective (TRL experimental PPO), then smoke tasks poe train-reward / poe train-ppo after an SFT warm-start. RLOO is a fallback only if experimental PPO proves too flaky. Phase 7 is VLM SFT separately.
🧑💻🤝 Contributing to SlickTune 🧩
You can find the details of the development process in our Contributing guidelines. We strongly believe that reading and following these guidelines will help us make the contribution process easy and effective for everyone involved 🚀🌙.
Special thanks to all of our amazing contributors 👇
❓ 🆘 📲 Need Help?
Please join our Slack Channel to interact directly with the core team and our small community. This is a good place to discuss your questions and ideas or in general ask for help 👨👩👧 👫 👨👩👦.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file slicktune-0.5.0.tar.gz.
File metadata
- Download URL: slicktune-0.5.0.tar.gz
- Upload date:
- Size: 371.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.9.13 {"installer":{"name":"uv","version":"0.9.13"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d0f8a86ad39af999655be6d8ca82047f387fdc50d7cb0ad121469a57e196b1f9
|
|
| MD5 |
1ce5709c36b387e260679516f52ef1ee
|
|
| BLAKE2b-256 |
d792a291c547f6ac9aff88d97cb4d9d5d1941083a71fa279b3132222f8740e46
|
File details
Details for the file slicktune-0.5.0-py3-none-any.whl.
File metadata
- Download URL: slicktune-0.5.0-py3-none-any.whl
- Upload date:
- Size: 39.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.9.13 {"installer":{"name":"uv","version":"0.9.13"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3943dd67142f7a4b79246298874c3e21f74c6f4970e43ad46a1717dcb8a6084a
|
|
| MD5 |
ea094e69d239abedaabb149df17bbfb5
|
|
| BLAKE2b-256 |
a614779d3fa87189288acca30add5d0c0eac2c333bcad3ed900d84eaa97da6c3
|