Skip to main content

build codecov downloads license pypi_version python_version slack_invite twitter_url

slick-tune logo

SlickTune 🧩: Composable LLM Fine-Tuning by SlickML🧞

Explore Releases 🟣 Become a Contributor 🟣 PyPI 🟣 Join our Slack 🟣 Tweet Us

🧠 Philosophy

SlickTune 🧩 is a small, composable toolkit for teaching LLMs new facts and behaviors with Transformers + PEFT + TRL. LoRA / QLoRA are PEFT adapters; full FT updates every weight. The goal is the same SlickML spirit: prototype fast 🏎, keep axes orthogonal, and measure whether the model actually learned your facts 🔎.

New to fine-tuning? Start here → Fine-Tuning LLMs: A Visual Guide — pre-training vs prompting vs FT, Full / LoRA / DoRA / AdaLoRA / QLoRA with diagrams, multi-adapter merge (TIES / DARE), how to choose a strategy, and probes & holdout PPL.

Fine-tuning is an orthogonal stack — swap any axis without rewriting the others:

model  ×  strategy  ×  objective  ×  data  ×  metrics

🧩 Abstractions

flowchart TB
  subgraph inputs [Inputs]
    modelId[model_id]
    dataJsonl[data JSONL]
  end

  subgraph axes [Composable axes]
    strategyNode["Strategy: LoRA / DoRA / AdaLoRA / QLoRA / Full"]
    objectiveNode["Objective: SFT / DPO / ORPO / KTO / GRPO"]
  end

  subgraph core [Tuner fit]
    tuner[Tuner]
    loadStep[load model and tokenizer]
    applyStep[strategy.apply]
    trainStep[TRL trainer]
    metricsStep[MetricsTracker]
  end

  subgraph outputs [Outputs]
    checkpoint[adapter or checkpoint]
    metricsFile[metrics.json]
    probeRate[probe pass rate]
  end

  modelId --> tuner
  dataJsonl --> tuner
  strategyNode --> tuner
  objectiveNode --> tuner
  tuner --> loadStep --> applyStep --> trainStep --> metricsStep
  trainStep --> checkpoint
  metricsStep --> metricsFile
  checkpoint --> probeRate
Axis Responsibility Shipped (phases 0–5)
Strategy How weights change (PEFT vs full) LoRA / DoRA / AdaLoRA / QLoRA / Full
Objective What is optimized / data contract SFT / DPO / ORPO / KTO / GRPO
Data Examples → chat, prefs, or rewards train + holdout + prefs/KTO/GRPO JSONL (about_amir*.jsonl)
Metrics Comparable run stats MetricsTracker (+ holdout PPL, judge score)
Merge Combine / bake adapters TIES / DARE / linear + bake_adapter
Eval Holdout + judges slicktune eval, SubstringJudge, LLMJudge
Probe Did the model learn your facts? slicktune probe

📌 Quick Start

from slicktune import LoRAStrategy, SFTObjective, Tuner

Tuner(
    model_id="HuggingFaceTB/SmolLM2-135M-Instruct",
    strategy=LoRAStrategy(r=8),
    objective=SFTObjective(),
    output_dir="outputs/sft_lora",
    eval_data="examples/data/about_amir.eval.jsonl",
).fit("examples/data/about_amir.jsonl")

👤 Personal “about me” loop (recommended)

  1. Edit examples/data/about_amir.jsonl with facts about you (or keep the SlickML starter facts) ✍️.
  2. Edit examples/data/about_amir.eval.jsonl with held-out paraphrases (same topics, not copied from train) for perplexity 📉.
  3. Edit examples/data/about_amir.probes.jsonl with questions and a must_contain substring that should appear after training 🎯.
  4. Train a strategy on a tiny instruct model 🧪.
  5. Probe the checkpoint — pass rate shows whether fine-tuning stuck ✅.
before FT  →  model guesses / hallucinates about you
after FT   →  probe answers contain your facts

🛠 Installation

Install Python >=3.10,<3.14 and uv, then simply run 🏃‍♀️:

uv sync --locked --all-extras --all-groups

QLoRA (CUDA + bitsandbytes only) 🔥:

uv sync --extra qlora

Task runner is Poe the Poet (same idea as slick-ml, with uv instead of Poetry). Install the CLI once 🏃‍♀️:

uv tool install poethepoet
poe greet

Developer workflow (format / check / test) lives in CONTRIBUTING.md 🧑‍💻🤝.

🚂 Train each strategy

Default demo model: HuggingFaceTB/SmolLM2-135M-Instruct (small enough for laptop smoke tests) 💻.

🟢 LoRA + SFT (default — works on Mac MPS / CPU / CUDA)

uv run slicktune train \
  --strategy lora \
  --data examples/data/about_amir.jsonl \
  --eval-data examples/data/about_amir.eval.jsonl \
  --output outputs/sft_lora \
  --epochs 20

uv run slicktune probe \
  --model-dir outputs/sft_lora \
  --probes examples/data/about_amir.probes.jsonl

Or: poe train-lora / poe probe-lora / poe eval-lora / uv run python examples/run_sft_lora.py

🟣 DoRA + SFT

uv run slicktune train \
  --strategy dora \
  --data examples/data/about_amir.jsonl \
  --eval-data examples/data/about_amir.eval.jsonl \
  --output outputs/sft_dora
# or: uv run python examples/run_sft_dora.py

🟤 AdaLoRA + SFT

uv run slicktune train \
  --strategy adalora \
  --data examples/data/about_amir.jsonl \
  --eval-data examples/data/about_amir.eval.jsonl \
  --output outputs/sft_adalora
# or: uv run python examples/run_sft_adalora.py

🟡 LoRA + DPO (preference pairs)

uv run slicktune train \
  --strategy lora \
  --objective dpo \
  --data examples/data/about_amir.prefs.jsonl \
  --output outputs/dpo_lora \
  --epochs 10

# or: poe train-dpo / uv run python examples/run_dpo_lora.py

🟢 LoRA + KTO (unpaired labels)

uv run slicktune train \
  --strategy lora \
  --objective kto \
  --data examples/data/about_amir.kto.jsonl \
  --output outputs/kto_lora \
  --epochs 10

# or: poe train-kto / uv run python examples/run_kto_lora.py

ORPO: --objective orpo with the same prefs JSONL as DPO (TRL experimental).

🔴 LoRA + GRPO (verifiable substring rewards)

GRPO samples multiple completions per prompt and scores them with a verifiable must_contain reward (exact match = 1.0, else keyword-overlap fraction). On a cold tiny base model rewards stay ~0 so GRPO cannot learn — warm-start with SFT first. The CLI train command has no --adapter-path; use the smoke example (or Tuner(..., adapter_path=...) in Python):

# SFT warm-start → GRPO (writes outputs/grpo_lora_sft then outputs/grpo_lora)
poe train-grpo
# or: uv run python examples/run_grpo_lora.py

🟣 Merge adapters (TIES / DARE)

Combine multiple PEFT adapters on the same base (TIES, DARE, linear, …), or bake into full weights for serving engines that want a single checkpoint. Smoke demo trains two tiny adapters then merges them:

poe merge-ties
# or: uv run python examples/run_merge_ties.py
# → outputs/merge_a_lora + outputs/merge_b_lora → outputs/merged_ties
uv run slicktune merge \
  --model HuggingFaceTB/SmolLM2-135M-Instruct \
  --adapter outputs/merge_a_lora \
  --adapter outputs/merge_b_lora:0.5 \
  --method ties \
  --density 0.5 \
  --output outputs/merged_ties

# bake into full weights: add --bake
# alternate: merge any two trained adapters, e.g. outputs/sft_lora + outputs/dpo_lora:0.5
from slicktune import AdapterRef, merge_adapters

merge_adapters(
    model_id="HuggingFaceTB/SmolLM2-135M-Instruct",
    adapters=[
        AdapterRef(path="outputs/merge_a_lora", name="a", weight=1.0),
        AdapterRef(path="outputs/merge_b_lora", name="b", weight=0.5),
    ],
    output_dir="outputs/merged_ties",
    method="ties",
    density=0.5,
)

🔎 Eval harness (holdout PPL + judges)

uv run slicktune eval \
  --model-dir outputs/sft_lora \
  --eval-data examples/data/about_amir.eval.jsonl \
  --probes examples/data/about_amir.probes.jsonl \
  --judge substring

--eval-data should be a holdout SFT JSONL (not the training file). The shipped about_amir.eval.jsonl paraphrases the same topics for holdout perplexity.

Use --judge llm to score generations with an LLM rubric (0–10 → normalized). On the tiny demo model, prefer --judge substring: the same 135M checkpoint is a weak judge and will under-score even when answers are correct.

🔵 QLoRA + SFT (CUDA required)

uv sync --extra qlora
uv run python examples/run_sft_qlora.py

On Apple Silicon, use LoRA instead — bitsandbytes 4-bit needs CUDA 🍎.

🟠 Full fine-tuning + SFT

uv run python examples/run_sft_full.py

Heavier on memory; prefer LoRA for iteration 💾.

📦 Data formats

SFT JSONL (any of these per line) 📝:

{"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}
{"prompt":"...","response":"..."}
{"instruction":"...","input":"...","output":"..."}

Probe JSONL 🕵️:

{"prompt":"Who is Amirhessam Tahmassebi?","must_contain":"SlickML"}

Holdout eval JSONL (same SFT shapes as train; keep examples out of the train file) 📉:

{"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}

Ship example: examples/data/about_amir.eval.jsonl.

Preference JSONL (DPO / ORPO) ⚖️:

{"prompt":"...","chosen":"...","rejected":"..."}

KTO JSONL (unpaired labels) ✅❌:

{"prompt":"...","completion":"...","label":true}

GRPO JSONL (verifiable RL; solution is accepted as an alias for must_contain) 🎯:

{"prompt":"Who is Amirhessam Tahmassebi?","must_contain":"founder of SlickML"}

🗺 Roadmap

Phase Scope
0–1 Skeleton, SFT + LoRA/QLoRA/full, metrics, personal probe loop
2 (done) DoRA / AdaLoRA, holdout PPL + substring/LLM judges
3 (done) DPO / ORPO / KTO
4 (done) GRPO / verifiable RL
5 (done) Merge (TIES/DARE), multi-adapter
6 (now) Classic RLHF: reward model + PPO (SFT → RM → online PPO)
7 Optional multimodal PEFT (VLM SFT + LoRA)

Planned (not implemented yet): Phase 6 adds RewardObjective + PPOObjective (TRL experimental PPO), then smoke tasks poe train-reward / poe train-ppo after an SFT warm-start. RLOO is a fallback only if experimental PPO proves too flaky. Phase 7 is VLM SFT separately.

🧑‍💻🤝 Contributing to SlickTune 🧩

You can find the details of the development process in our Contributing guidelines. We strongly believe that reading and following these guidelines will help us make the contribution process easy and effective for everyone involved 🚀🌙.

Special thanks to all of our amazing contributors 👇

Repobeats analytics image

❓ 🆘 📲 Need Help?

Please join our Slack Channel to interact directly with the core team and our small community. This is a good place to discuss your questions and ideas or in general ask for help 👨‍👩‍👧 👫 👨‍👩‍👦.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

slicktune-0.5.0.tar.gz (371.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

slicktune-0.5.0-py3-none-any.whl (39.9 kB view details)

Uploaded Python 3

File details

Details for the file slicktune-0.5.0.tar.gz.

File metadata

  • Download URL: slicktune-0.5.0.tar.gz
  • Upload date:
  • Size: 371.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.13 {"installer":{"name":"uv","version":"0.9.13"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for slicktune-0.5.0.tar.gz
Algorithm Hash digest
SHA256 d0f8a86ad39af999655be6d8ca82047f387fdc50d7cb0ad121469a57e196b1f9
MD5 1ce5709c36b387e260679516f52ef1ee
BLAKE2b-256 d792a291c547f6ac9aff88d97cb4d9d5d1941083a71fa279b3132222f8740e46

See more details on using hashes here.

File details

Details for the file slicktune-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: slicktune-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 39.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.13 {"installer":{"name":"uv","version":"0.9.13"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for slicktune-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3943dd67142f7a4b79246298874c3e21f74c6f4970e43ad46a1717dcb8a6084a
MD5 ea094e69d239abedaabb149df17bbfb5
BLAKE2b-256 a614779d3fa87189288acca30add5d0c0eac2c333bcad3ed900d84eaa97da6c3

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page