Skip to main content

Agent + CLI that simplifies fine-tuning with Unsloth, adding complementary actions so an agent can fine-tune models more easily.

Project description

unsloth-cli

Agent + CLI that simplifies fine-tuning with Unsloth, adding complementary actions so an agent can fine-tune models more easily.

What you get

  • An agent-first CLI cited from teken (afi-cli) — the runtime package has no third-party dependencies.
  • A mesh identityculture.yaml (suffix + backend) and the matching prompt file (CLAUDE.md for backend: claude).
  • The canonical guildmaster skill kit (11 skills) under .claude/skills/, vendored cite-don't-import. See docs/skill-sources.md.
  • A build + deploy baseline — pytest, lint, the agent-first rubric gate, and PyPI Trusted Publishing wired into GitHub Actions.

Quickstart

uv sync
uv run pytest -n auto                 # run the test suite
uv run sloth whoami                   # identity from culture.yaml
uv run sloth learn                    # self-teaching prompt (add --json)
uv run teken cli doctor . --strict    # the agent-first rubric gate CI runs

The installed console script is sloth (the dist name is unsloth-cli); run sloth <verb> or python -m sloth <verb>. The CLI prints unsloth-cli in its help/explain text because that is the argparse program name.

CLI

Verb What it does
whoami Report this agent's nick, version, backend, and model from culture.yaml.
learn Print a structured self-teaching prompt.
explain <path> Markdown docs for any noun/verb path.
overview Read-only descriptive snapshot of the agent.
doctor Check the agent-identity invariants (prompt-file-present, backend-consistency).
cli overview Describe the CLI surface itself.

Every command supports --json. Results go to stdout, errors/diagnostics to stderr (never mixed). Exit codes: 0 success, 1 user error, 2 environment error, 3+ reserved.

Fine-tuning

unsloth-cli ships three flat verbs for LoRA/QLoRA adapter tuning of Qwen models, plus a /finetune skill that drives the full loop. The Unsloth/PyTorch stack arrives with uv tool install unsloth-cli; the introspection verbs (whoami, learn, explain, etc.) stay fast via lazy imports — torch is never loaded at module import time.

Out of scope

Full fine-tuning of large dense models is not supported. The CLI targets LoRA and QLoRA adapters on small-to-medium Qwen models (Qwen 3.x 4B / 9B and comparable adapter-class targets). Pointing sloth train at a large dense full-fine-tune target emits an explicit warning and refuses or downgrades to adapter-only — it does not attempt the job silently.

Commands

Verb What it does
sloth train Validate JSONL dataset → run LoRA/QLoRA adapter job → write run metadata
sloth eval Run an adapter against a small local eval suite (no network)
sloth export Convert an adapter to safetensors (servable by lobes, runnable by colleague)

The /finetune skill drives the full loop non-interactively: validate dataset → sloth trainsloth evalsloth export.

Every verb supports --json and routes errors through error: / hint: on stderr.

Dataset schemas

Two JSONL schemas are supported. Validation runs before spending any GPU time; malformed lines are reported with the offending line number and a remediation hint.

Chat format — for instruction-following and conversational behavior:

{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}

Task format — for structured input/output tasks:

{"task": "write-issue", "input": "...", "expected_output": "..."}

Run config (TOML) and Spark-friendly defaults

Training runs are driven by a TOML config file. Omitted optional keys fall back to Spark-friendly defaults tuned for small-GPU (single-card Spark) operation.

[run]
model   = "unsloth/Qwen3-4B"      # supported: Qwen3 4B / 9B adapter-class targets
method  = "qlora"                 # "lora" or "qlora" — the only supported methods (default: qlora)
dataset = "data/train.jsonl"
output  = "adapters/my-lora"

[hyperparameters]
lora_r        = 16                # LoRA rank          (default: 16)
lora_alpha    = 16                # LoRA alpha scaling (default: 16)
lora_dropout  = 0.0               # default: 0.0
learning_rate = 2e-4              # default: 2e-4
max_seq_len   = 2048              # default: 2048
batch_size    = 2                 # default: 2  (Spark-friendly: keeps VRAM low)
grad_accum    = 4                 # default: 4
max_steps     = 60                # default: 60 (quick smoke-run; raise for production)
seed          = 3407              # default: 3407
load_in_4bit  = true              # default: true (required for qlora)

A metadata file is written next to the adapter output recording model, method, dataset SHA-256 and line count, hyperparameters, and an ISO-8601 timestamp. Re-running the same config file and dataset reproduces the same training setup.

What belongs in fine-tuning vs. memory / RAG

This is a design rule, not a footnote. The fine-tune/RAG boundary decides where a capability lives in the mesh.

Fine-tune stores stable behavior and reflexes — things that should be baked into how the model responds, not looked up on every call:

  • CLI-contract discipline (error/hint format, exit-code policy, stream split)
  • AgentCulture / CULTURE.DEV terminology and patterns
  • Agent-first habits (prefer action verbs, emit structured --json, route errors correctly)
  • Issue-writing format and AgentCulture PR/review norms
  • Teacher behavior for learn and explain responses

Memory / RAG stores changing facts — things that vary per session, user, or deployment and would become stale if baked into weights:

  • Current project state, open issues, branch status, recent commits
  • Secrets, tokens, credentials, or any per-deployment configuration
  • User-specific preferences or operator-specific memory
  • Facts better served by retrieval (live documentation, changelogs, external APIs)

Decision rule for contributors: "Would this still be correct six months from now on any deployment of the mesh?" If yes, consider fine-tuning. If it changes over time or is per-user, use memory / RAG.

Role-specific adapters

The design targets small, role-specific adapters rather than one large mixed blob. Example adapter names that map to discrete behaviors:

  • culture-contract-lora — CLI-contract discipline and AgentCulture norms
  • agentculture-cli-teacher-lora — teacher behavior for learn / explain
  • repo-maintainer-lora — issue-writing format and PR review norms
  • tool-router-lora — tool selection and routing decisions
  • agent-first-coach-lora — agent-first habits and patterns

The resulting adapters are written in standard PEFT / safetensors layout so lobes can serve them and colleague can run them as model backends.

Make it your own

  1. Rename the package sloth/ and the unsloth-cli CLI/dist name throughout pyproject.toml, the package, tests/, sonar-project.properties, and this README.md. The name is hard-coded in ~100 places, so list every occurrence first — see the git grep discovery command in CLAUDE.md, the authoritative rename procedure.
  2. Edit culture.yaml with your suffix and backend.
  3. Rewrite CLAUDE.md for your agent and run /init.
  4. Re-vendor only the skills you need from guildmaster (see docs/skill-sources.md).

See CLAUDE.md for the full conventions (version-bump-every-PR, the cicd PR lane, deploy setup).

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

unsloth_cli-0.4.0.tar.gz (312.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

unsloth_cli-0.4.0-py3-none-any.whl (45.8 kB view details)

Uploaded Python 3

File details

Details for the file unsloth_cli-0.4.0.tar.gz.

File metadata

  • Download URL: unsloth_cli-0.4.0.tar.gz
  • Upload date:
  • Size: 312.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.24 {"installer":{"name":"uv","version":"0.11.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for unsloth_cli-0.4.0.tar.gz
Algorithm Hash digest
SHA256 673b30992690b6c9bab0a3f140a1a0d66a189f0e5c7239ceffe3283f743970cf
MD5 41304c6436e3f13ac72ca1c8682cf6a8
BLAKE2b-256 1ac8a4057c8dbe07007c8e9be409d2ab3c747ee2ab20ac99c671ab752474e347

See more details on using hashes here.

File details

Details for the file unsloth_cli-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: unsloth_cli-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 45.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.24 {"installer":{"name":"uv","version":"0.11.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for unsloth_cli-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 5e89e18cecb844177cbb5c5e0980dc6aecfba86ef0d98456232f05680479588a
MD5 9328c0a7a7b169785dd7af653be953bf
BLAKE2b-256 d75be53fbfb83ccd4e45819634ef1233ebf39c3e76621534c6cfea547a425df8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page