Agent + CLI that simplifies fine-tuning with Unsloth, adding complementary actions so an agent can fine-tune models more easily.
Project description
unsloth-cli
Agent + CLI that simplifies fine-tuning with Unsloth, adding complementary actions so an agent can fine-tune models more easily.
What you get
- An agent-first CLI cited from teken
(
afi-cli) — the runtime package has no third-party dependencies. - A mesh identity —
culture.yaml(suffix+backend) and the matching prompt file (CLAUDE.mdforbackend: claude). - The canonical guildmaster skill kit (11 skills) under
.claude/skills/, vendored cite-don't-import. Seedocs/skill-sources.md. - A build + deploy baseline — pytest, lint, the agent-first rubric gate, and PyPI Trusted Publishing wired into GitHub Actions.
Quickstart
uv sync
uv run pytest -n auto # run the test suite
uv run sloth whoami # identity from culture.yaml
uv run sloth learn # self-teaching prompt (add --json)
uv run teken cli doctor . --strict # the agent-first rubric gate CI runs
The installed console script is sloth (the dist name is unsloth-cli); run
sloth <verb> or python -m sloth <verb>. The CLI prints unsloth-cli in its
help/explain text because that is the argparse program name.
CLI
| Verb | What it does |
|---|---|
whoami |
Report this agent's nick, version, backend, and model from culture.yaml. |
learn |
Print a structured self-teaching prompt. |
explain <path> |
Markdown docs for any noun/verb path. |
overview |
Read-only descriptive snapshot of the agent. |
doctor |
Check the agent-identity invariants (prompt-file-present, backend-consistency). |
cli overview |
Describe the CLI surface itself. |
Every command supports --json. Results go to stdout, errors/diagnostics to
stderr (never mixed). Exit codes: 0 success, 1 user error, 2 environment
error, 3+ reserved.
Fine-tuning
unsloth-cli ships three flat verbs for LoRA/QLoRA adapter tuning of Qwen models,
plus a /finetune skill that drives the full loop. The Unsloth/PyTorch stack
arrives with uv tool install unsloth-cli; the introspection verbs (whoami,
learn, explain, etc.) stay fast via lazy imports — torch is never loaded at
module import time.
Out of scope
Full fine-tuning of large dense models is not supported. The CLI targets
LoRA and QLoRA adapters on small-to-medium Qwen models (Qwen 3.x 4B / 9B and
comparable adapter-class targets). Pointing sloth train at a large dense
full-fine-tune target emits an explicit warning and refuses or downgrades to
adapter-only — it does not attempt the job silently.
Commands
| Verb | What it does |
|---|---|
sloth train |
Validate JSONL dataset → run LoRA/QLoRA adapter job → write run metadata |
sloth eval |
Run an adapter against a small local eval suite (no network) |
sloth export |
Convert an adapter to safetensors (servable by lobes, runnable by colleague) |
The /finetune skill drives the full loop non-interactively:
validate dataset → sloth train → sloth eval → sloth export.
Every verb supports --json and routes errors through error: / hint: on stderr.
Dataset schemas
Two JSONL schemas are supported. Validation runs before spending any GPU time; malformed lines are reported with the offending line number and a remediation hint.
Chat format — for instruction-following and conversational behavior:
{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Task format — for structured input/output tasks:
{"task": "write-issue", "input": "...", "expected_output": "..."}
Run config (TOML) and Spark-friendly defaults
Training runs are driven by a TOML config file. Omitted optional keys fall back to Spark-friendly defaults tuned for small-GPU (single-card Spark) operation.
[run]
model = "unsloth/Qwen3-4B" # supported: Qwen3 4B / 9B adapter-class targets
method = "qlora" # "lora" or "qlora" — the only supported methods (default: qlora)
dataset = "data/train.jsonl"
output = "adapters/my-lora"
[hyperparameters]
lora_r = 16 # LoRA rank (default: 16)
lora_alpha = 16 # LoRA alpha scaling (default: 16)
lora_dropout = 0.0 # default: 0.0
learning_rate = 2e-4 # default: 2e-4
max_seq_len = 2048 # default: 2048
batch_size = 2 # default: 2 (Spark-friendly: keeps VRAM low)
grad_accum = 4 # default: 4
max_steps = 60 # default: 60 (quick smoke-run; raise for production)
seed = 3407 # default: 3407
load_in_4bit = true # default: true (required for qlora)
A metadata file is written next to the adapter output recording model, method, dataset SHA-256 and line count, hyperparameters, and an ISO-8601 timestamp. Re-running the same config file and dataset reproduces the same training setup.
What belongs in fine-tuning vs. memory / RAG
This is a design rule, not a footnote. The fine-tune/RAG boundary decides where a capability lives in the mesh.
Fine-tune stores stable behavior and reflexes — things that should be baked into how the model responds, not looked up on every call:
- CLI-contract discipline (error/hint format, exit-code policy, stream split)
- AgentCulture / CULTURE.DEV terminology and patterns
- Agent-first habits (prefer action verbs, emit structured
--json, route errors correctly) - Issue-writing format and AgentCulture PR/review norms
- Teacher behavior for
learnandexplainresponses
Memory / RAG stores changing facts — things that vary per session, user, or deployment and would become stale if baked into weights:
- Current project state, open issues, branch status, recent commits
- Secrets, tokens, credentials, or any per-deployment configuration
- User-specific preferences or operator-specific memory
- Facts better served by retrieval (live documentation, changelogs, external APIs)
Decision rule for contributors: "Would this still be correct six months from now on any deployment of the mesh?" If yes, consider fine-tuning. If it changes over time or is per-user, use memory / RAG.
Role-specific adapters
The design targets small, role-specific adapters rather than one large mixed blob. Example adapter names that map to discrete behaviors:
culture-contract-lora— CLI-contract discipline and AgentCulture normsagentculture-cli-teacher-lora— teacher behavior forlearn/explainrepo-maintainer-lora— issue-writing format and PR review normstool-router-lora— tool selection and routing decisionsagent-first-coach-lora— agent-first habits and patterns
The resulting adapters are written in standard PEFT / safetensors layout so lobes can serve them and colleague can run them as model backends.
Make it your own
- Rename the package
sloth/and theunsloth-cliCLI/dist name throughoutpyproject.toml, the package,tests/,sonar-project.properties, and thisREADME.md. The name is hard-coded in ~100 places, so list every occurrence first — see thegit grepdiscovery command inCLAUDE.md, the authoritative rename procedure. - Edit
culture.yamlwith yoursuffixandbackend. - Rewrite
CLAUDE.mdfor your agent and run/init. - Re-vendor only the skills you need from guildmaster (see
docs/skill-sources.md).
See CLAUDE.md for the full conventions (version-bump-every-PR,
the cicd PR lane, deploy setup).
License
MIT — see LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file unsloth_cli-0.4.0.tar.gz.
File metadata
- Download URL: unsloth_cli-0.4.0.tar.gz
- Upload date:
- Size: 312.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.24 {"installer":{"name":"uv","version":"0.11.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
673b30992690b6c9bab0a3f140a1a0d66a189f0e5c7239ceffe3283f743970cf
|
|
| MD5 |
41304c6436e3f13ac72ca1c8682cf6a8
|
|
| BLAKE2b-256 |
1ac8a4057c8dbe07007c8e9be409d2ab3c747ee2ab20ac99c671ab752474e347
|
File details
Details for the file unsloth_cli-0.4.0-py3-none-any.whl.
File metadata
- Download URL: unsloth_cli-0.4.0-py3-none-any.whl
- Upload date:
- Size: 45.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.24 {"installer":{"name":"uv","version":"0.11.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5e89e18cecb844177cbb5c5e0980dc6aecfba86ef0d98456232f05680479588a
|
|
| MD5 |
9328c0a7a7b169785dd7af653be953bf
|
|
| BLAKE2b-256 |
d75be53fbfb83ccd4e45819634ef1233ebf39c3e76621534c6cfea547a425df8
|