The Classifier
A general-purpose natural language classifier with domain-specific LoRA adapters.
Run locally on CPU or NVIDIA GPU, serve the same /v1/classify request format as
the hosted service, or fine-tune an adapter on your own labelled examples.
Release status
This is the initial 0.1.0 package. The configured model repository is
TrueState/theclassifier. The checkpoints are staged privately; authenticate with
Hugging Face and obtain repository access before using named-version downloads.
A local checkpoint also works.
Model weights are not included in the Python package. This code is Apache-2.0;
checkpoint and dataset licences are separate.
Install
From this checkout, until the PyPI release is published:
pip install './python[serve]'
After release:
pip install 'theclassifier[serve]'
Python 3.10+ is required. Install a PyTorch build suitable for your platform first
if needed; CUDA support depends on that build and your NVIDIA driver. CPU uses
float32; CUDA uses bfloat16 when supported, otherwise float32. No FlashAttention
extension is required. auto selects CUDA when available and CPU otherwise.
Apple MPS is not currently supported.
Classify locally
from theclassifier import Classifier
classifier = Classifier() # general; downloads once, then uses the local cache
result = classifier.classify(
'Please close my account.',
choices=['billing', 'cancellation', 'none'],
question='What is the customer asking for?',
)
print(result['results'][0]['selected'])
Use Classifier(checkpoint='./checkpoints/general', device='cpu') to load an
existing local checkpoint. A checkpoint contains adapter/, tokenizer/ and
head.safetensors (or the existing head.pt format). The pretrained backbone is
resolved from the adapter configuration. A local adapter can still require a
backbone download if the backbone is not already cached.
The full question format supports multiple decisions and optional criteria:
result = classifier.classify(
'Please close my account.',
questions=[{
'question': 'What is the customer asking for?',
'choices': [
{'choice': 'cancellation', 'criteria': 'Requests to close an account.'},
{'choice': 'none', 'criteria': 'None of the other choices apply.'},
],
}],
version='intent',
)
Versions: general, intent, sentiment, triage, cause, codify,
complaints, stars. Selecting a new version checks the Hugging Face cache,
downloads missing files, then loads the adapter and its scoring head onto the
shared backbone. Selection and inference are serialized to keep the adapter and
head paired. A failed download raises an error; it never silently falls back to
another model.
classifier.load_adapter('custom', checkpoint='./checkpoints/my-adapter')
result = classifier.classify('example', choices=['yes', 'no'], version='custom')
Use model_repo=, revision= and cache_dir= or the
THECLASSIFIER_MODEL_REPO environment variable to override download settings.
Pin revision to a commit for reproducibility. offline=True disables network
loading, including the base model. Pre-download adapters with:
theclassifier download --version general
theclassifier download --version intent
These commands download adapters and tokenizers; initialize the classifier once
online to cache the backbone as well. The repository layout is
<version>/adapter/, <version>/tokenizer/, <version>/head.safetensors.
Default maximum input length is 2,048 tokens per candidate; longer inputs retain
the first token and tail, matching the existing prompt format. Set max_length
explicitly for longer records. batch_size controls inference candidate chunks.
This initial portable runtime uses direct candidate scoring, not shared-prefix
KV caching. It has not been benchmarked against the optimized production server.
Serve on CPU or GPU
theclassifier serve --checkpoint ./checkpoints/general --device cpu
# On a CUDA machine:
API_KEY=your-serving-key theclassifier serve --device cuda --host 0.0.0.0
Default binding is 127.0.0.1:8090. Binding outside localhost requires API_KEY.
Use one process per GPU; run behind a TLS proxy for remote clients. This standalone
server provides inference and authentication, not the hosted service's billing,
quotas or rate limiting. Set resource limits in your deployment.
curl http://127.0.0.1:8090/v1/classify \
-H 'Content-Type: application/json' \
-d '{"state":"Close my account","questions":[{"question":"Intent?","choices":[{"choice":"cancel"},{"choice":"none"}]}]}'
Include Authorization: Bearer YOUR_KEY when API_KEY is configured.
Responses contain results (with question, selected, detailed_results),
version, input_tokens and latency_ms.
Fine-tune with LoRA
Each JSONL line is one complete choice menu, with an exact matching answer label:
{"state":"Close my account","question":"Intent?","choices":[{"choice":"cancel","criteria":"Requests to close an account"},{"choice":"none"}],"answer":"cancel"}
# Continue the general checkpoint to create a specialist:
theclassifier finetune --version general --data train.jsonl \
--output checkpoints/my-adapter --device cuda
# Train an adapter and scoring head from the base model:
theclassifier finetune --base-model Qwen/Qwen3-1.7B-Base \
--data train.jsonl --output checkpoints/from-base --rank 16 --alpha 16
Or in Python:
from theclassifier import finetune
report = finetune(
'train.jsonl', 'checkpoints/my-adapter',
checkpoint='checkpoints/general', device='auto', epochs=1,
learning_rate=2e-4, grad_accum=16,
)
Training freezes the backbone, learns LoRA updates to q_proj and v_proj, and
trains a float32 scalar head using cross-entropy across each menu. Whole menus
are processed together, gradients accumulate between menus, and gradient
checkpointing reduces activation memory. Large choice menus may still require
smaller input lengths or more memory. CPU training is supported but slow.
New adapters default to rank 16, alpha 16 and dropout 0.05. Continuing an existing
adapter preserves its LoRA configuration. A new AdamW optimizer is created with
a constant learning rate; this is fine-tuning, not exact training-state resume.
The optimizer configuration differs from the original research runs. Output
folders must be empty. Saved files include the adapter, head, tokenizer,
classifier.json and training.json. Keep evaluation data separate and evaluate
before using a new checkpoint; the training loss is not a quality benchmark.
Development and release
cd python
pip install -e '.[serve,dev]'
pytest
python -m build
python -m twine check dist/*
# With PyPI credentials configured locally:
python -m twine upload dist/*
The tests build a tiny random Qwen model locally: no production weights, private
data or external inference service is needed. All six tests passed on the NVIDIA
workstation, including CUDA inference. Real v1.1 checkpoints were also smoke-tested
on CUDA, and the general model was checked on CPU against the original scoring
implementation. These checks establish execution compatibility, not accuracy.
Publishing the package does not upload model weights. See RELEASE.md for the
checkpoint and publication steps.
Metadata
Release files for theclassifier 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| theclassifier-0.1.0.tar.gz | 22.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| theclassifier-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 42.1 kB
Release files / theclassifier-0.1.0.tar.gz
| Download URL | theclassifier-0.1.0.tar.gz |
|---|---|
| Size | 22.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7d21dda2ec312a78d8bc9f7e1b43dfcf33b9da77479f5eeb8d0e77e1b7b96b77
|
|
BLAKE2b-256 checksum How to use checksums |
7c6a2f4a0e08aa61dd9f19a299125d4f557aef98c19651c65c4ef667a9747ea1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|
Release files / theclassifier-0.1.0-py3-none-any.whl
| Download URL | theclassifier-0.1.0-py3-none-any.whl |
|---|---|
| Size | 19.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
95dd10becbae622f1f690a70f94fe0e1b1751c230469cdbd668aac283d732261
|
|
BLAKE2b-256 checksum How to use checksums |
287bfdeece50494d58fa4dd045b51e9f91ffb11f2577e6ea4437e91b06730834
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|