Skip to main content

build codecov downloads license pypi_version python_version slack_invite twitter_url

slick-tune logo

SlickTune 🧩: Composable LLM Fine-Tuning by SlickML🧞

Explore Releases 🟣 Become a Contributor 🟣 PyPI 🟣 Join our Slack 🟣 Tweet Us

🧠 Philosophy

SlickTune 🧩 is a small, composable toolkit for teaching LLMs new facts and behaviors with Transformers + PEFT + TRL. LoRA / QLoRA are PEFT adapters; full FT updates every weight. The goal is the same SlickML spirit: prototype fast 🏎, keep axes orthogonal, and measure whether the model actually learned your facts 🔎.

New to fine-tuning? Start here → Fine-Tuning LLMs: A Visual Guide — pre-training vs prompting vs FT, Full / LoRA / DoRA / AdaLoRA / QLoRA with diagrams, how to choose a strategy, and how probes & holdout perplexity tell you it worked.

Fine-tuning is an orthogonal stack — swap any axis without rewriting the others:

model  ×  strategy  ×  objective  ×  data  ×  metrics

🧩 Abstractions

flowchart TB
  subgraph inputs [Inputs]
    modelId[model_id]
    dataJsonl[data JSONL]
  end

  subgraph axes [Composable axes]
    strategyNode["Strategy: LoRA / DoRA / AdaLoRA / QLoRA / Full"]
    objectiveNode["Objective: SFT then DPO / GRPO"]
  end

  subgraph core [Tuner fit]
    tuner[Tuner]
    loadStep[load model and tokenizer]
    applyStep[strategy.apply]
    trainStep[TRL trainer]
    metricsStep[MetricsTracker]
  end

  subgraph outputs [Outputs]
    checkpoint[adapter or checkpoint]
    metricsFile[metrics.json]
    probeRate[probe pass rate]
  end

  modelId --> tuner
  dataJsonl --> tuner
  strategyNode --> tuner
  objectiveNode --> tuner
  tuner --> loadStep --> applyStep --> trainStep --> metricsStep
  trainStep --> checkpoint
  metricsStep --> metricsFile
  checkpoint --> probeRate
Axis Responsibility Phase 4
Strategy How weights change (PEFT vs full) LoRA / DoRA / AdaLoRA / QLoRA / Full
Objective What is optimized / data contract SFT / DPO / ORPO / KTO / GRPO
Data Examples → chat, prefs, or rewards train + holdout + prefs/KTO/GRPO JSONL (about_amir*.jsonl)
Metrics Comparable run stats MetricsTracker (+ holdout PPL, judge score)
Eval Holdout + judges slicktune eval, SubstringJudge, LLMJudge
Probe Did the model learn your facts? slicktune probe

📌 Quick Start

from slicktune import LoRAStrategy, SFTObjective, Tuner

Tuner(
    model_id="HuggingFaceTB/SmolLM2-135M-Instruct",
    strategy=LoRAStrategy(r=8),
    objective=SFTObjective(),
    output_dir="outputs/sft_lora",
    eval_data="examples/data/about_amir.eval.jsonl",
).fit("examples/data/about_amir.jsonl")

👤 Personal “about me” loop (recommended)

  1. Edit examples/data/about_amir.jsonl with facts about you (or keep the SlickML starter facts) ✍️.
  2. Edit examples/data/about_amir.eval.jsonl with held-out paraphrases (same topics, not copied from train) for perplexity 📉.
  3. Edit examples/data/about_amir.probes.jsonl with questions and a must_contain substring that should appear after training 🎯.
  4. Train a strategy on a tiny instruct model 🧪.
  5. Probe the checkpoint — pass rate shows whether fine-tuning stuck ✅.
before FT  →  model guesses / hallucinates about you
after FT   →  probe answers contain your facts

🛠 Installation

Install Python >=3.10,<3.14 and uv, then simply run 🏃‍♀️:

uv sync --locked --all-extras --all-groups

QLoRA (CUDA + bitsandbytes only) 🔥:

uv sync --extra qlora

Task runner is Poe the Poet (same idea as slick-ml, with uv instead of Poetry). Install the CLI once 🏃‍♀️:

uv tool install poethepoet
poe greet

Developer workflow (format / check / test) lives in CONTRIBUTING.md 🧑‍💻🤝.

🚂 Train each strategy

Default demo model: HuggingFaceTB/SmolLM2-135M-Instruct (small enough for laptop smoke tests) 💻.

🟢 LoRA + SFT (default — works on Mac MPS / CPU / CUDA)

uv run slicktune train \
  --strategy lora \
  --data examples/data/about_amir.jsonl \
  --eval-data examples/data/about_amir.eval.jsonl \
  --output outputs/sft_lora \
  --epochs 20

uv run slicktune probe \
  --model-dir outputs/sft_lora \
  --probes examples/data/about_amir.probes.jsonl

Or: poe train-lora / poe probe-lora / poe eval-lora / uv run python examples/run_sft_lora.py

🟣 DoRA + SFT

uv run slicktune train \
  --strategy dora \
  --data examples/data/about_amir.jsonl \
  --eval-data examples/data/about_amir.eval.jsonl \
  --output outputs/sft_dora
# or: uv run python examples/run_sft_dora.py

🟤 AdaLoRA + SFT

uv run slicktune train \
  --strategy adalora \
  --data examples/data/about_amir.jsonl \
  --eval-data examples/data/about_amir.eval.jsonl \
  --output outputs/sft_adalora
# or: uv run python examples/run_sft_adalora.py

🟡 LoRA + DPO (preference pairs)

uv run slicktune train \
  --strategy lora \
  --objective dpo \
  --data examples/data/about_amir.prefs.jsonl \
  --output outputs/dpo_lora \
  --epochs 3

# or: poe train-dpo / uv run python examples/run_dpo_lora.py

🟢 LoRA + KTO (unpaired labels)

uv run slicktune train \
  --strategy lora \
  --objective kto \
  --data examples/data/about_amir.kto.jsonl \
  --output outputs/kto_lora \
  --epochs 3

# or: poe train-kto / uv run python examples/run_kto_lora.py

ORPO: --objective orpo with the same prefs JSONL as DPO (TRL experimental).

🔴 LoRA + GRPO (verifiable substring rewards)

uv run slicktune train \
  --strategy lora \
  --objective grpo \
  --data examples/data/about_amir.grpo.jsonl \
  --output outputs/grpo_lora \
  --epochs 3 \
  --num-generations 2 \
  --max-completion-length 64 \
  --beta 0.0

# or: poe train-grpo / uv run python examples/run_grpo_lora.py

GRPO samples multiple completions per prompt and scores them with a verifiable must_contain reward (exact match = 1.0, else keyword-overlap fraction). On a cold tiny base model rewards stay ~0 so GRPO cannot learn — warm-start with SFT first (poe train-grpo / examples/run_grpo_lora.py does this automatically).

🔎 Eval harness (holdout PPL + judges)

uv run slicktune eval \
  --model-dir outputs/sft_lora \
  --eval-data examples/data/about_amir.eval.jsonl \
  --probes examples/data/about_amir.probes.jsonl \
  --judge substring

--eval-data should be a holdout SFT JSONL (not the training file). The shipped about_amir.eval.jsonl paraphrases the same topics for holdout perplexity.

Use --judge llm to score generations with an LLM rubric (0–10 → normalized). On the tiny demo model, prefer --judge substring: the same 135M checkpoint is a weak judge and will under-score even when answers are correct.

🔵 QLoRA + SFT (CUDA required)

uv sync --extra qlora
uv run python examples/run_sft_qlora.py

On Apple Silicon, use LoRA instead — bitsandbytes 4-bit needs CUDA 🍎.

🟠 Full fine-tuning + SFT

uv run python examples/run_sft_full.py

Heavier on memory; prefer LoRA for iteration 💾.

📦 Data formats

SFT JSONL (any of these per line) 📝:

{"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}
{"prompt":"...","response":"..."}
{"instruction":"...","input":"...","output":"..."}

Probe JSONL 🕵️:

{"prompt":"Who is Amirhessam Tahmassebi?","must_contain":"SlickML"}

Holdout eval JSONL (same SFT shapes as train; keep examples out of the train file) 📉:

{"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}

Ship example: examples/data/about_amir.eval.jsonl.

Preference JSONL (DPO / ORPO) ⚖️:

{"prompt":"...","chosen":"...","rejected":"..."}

KTO JSONL (unpaired labels) ✅❌:

{"prompt":"...","completion":"...","label":true}

GRPO JSONL (verifiable RL; solution is accepted as an alias for must_contain) 🎯:

{"prompt":"Who is Amirhessam Tahmassebi?","must_contain":"founder of SlickML"}

🗺 Roadmap

Phase Scope
0–1 Skeleton, SFT + LoRA/QLoRA/full, metrics, personal probe loop
2 (done) DoRA / AdaLoRA, holdout PPL + substring/LLM judges
3 (done) DPO / ORPO / KTO
4 (done) GRPO / verifiable RL
5 (now) Merge (TIES/DARE), multi-adapter
6 Optional PPO / multimodal

🧑‍💻🤝 Contributing to SlickTune 🧩

You can find the details of the development process in our Contributing guidelines. We strongly believe that reading and following these guidelines will help us make the contribution process easy and effective for everyone involved 🚀🌙.

Special thanks to all of our amazing contributors 👇

Repobeats analytics image

❓ 🆘 📲 Need Help?

Please join our Slack Channel to interact directly with the core team and our small community. This is a good place to discuss your questions and ideas or in general ask for help 👨‍👩‍👧 👫 👨‍👩‍👦.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

slicktune-0.4.0.tar.gz (365.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

slicktune-0.4.0-py3-none-any.whl (34.4 kB view details)

Uploaded Python 3

File details

Details for the file slicktune-0.4.0.tar.gz.

File metadata

  • Download URL: slicktune-0.4.0.tar.gz
  • Upload date:
  • Size: 365.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.13 {"installer":{"name":"uv","version":"0.9.13"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for slicktune-0.4.0.tar.gz
Algorithm Hash digest
SHA256 8745108906d55292b201f437aa8d78b5527cfd7e45f1832eda02c0d43d8b3ab0
MD5 d223be09b7b72b9f63ca822d0c958e28
BLAKE2b-256 a72e0fc9c4b1c0743bd10c92d0a23864373ed8296a8d961950a6a5a94277ece5

See more details on using hashes here.

File details

Details for the file slicktune-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: slicktune-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 34.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.13 {"installer":{"name":"uv","version":"0.9.13"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for slicktune-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 aee09a25745e4a76944477b46758d9ed1db99937599d54b1b700156ff7f6f244
MD5 c2586eeebdad218798bb760c0b6718ed
BLAKE2b-256 aed2234e16b23cf5a31d1565aae31146db9b23a66a67a3250c8a0b1ad5d185c8

See more details on using hashes here.

Release history Release notifications | RSS feed

0.5.0

2 files

This release

0.4.0 This release

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page