Skip to main content

The Classifier

A general-purpose natural language classifier with domain-specific LoRA adapters. Run locally on CPU or NVIDIA GPU, serve the same /v1/classify request format as the hosted service, or fine-tune an adapter on your own labelled examples.

Release status

This is the initial 0.1.0 package. The configured model repository is TrueState/theclassifier. The checkpoints are staged privately; authenticate with Hugging Face and obtain repository access before using named-version downloads. A local checkpoint also works. Model weights are not included in the Python package. This code is Apache-2.0; checkpoint and dataset licences are separate.

Install

From this checkout, until the PyPI release is published:

pip install './python[serve]'

After release:

pip install 'theclassifier[serve]'

Python 3.10+ is required. Install a PyTorch build suitable for your platform first if needed; CUDA support depends on that build and your NVIDIA driver. CPU uses float32; CUDA uses bfloat16 when supported, otherwise float32. No FlashAttention extension is required. auto selects CUDA when available and CPU otherwise. Apple MPS is not currently supported.

Classify locally

from theclassifier import Classifier

classifier = Classifier()  # general; downloads once, then uses the local cache
result = classifier.classify(
    'Please close my account.',
    choices=['billing', 'cancellation', 'none'],
    question='What is the customer asking for?',
)
print(result['results'][0]['selected'])

Use Classifier(checkpoint='./checkpoints/general', device='cpu') to load an existing local checkpoint. A checkpoint contains adapter/, tokenizer/ and head.safetensors (or the existing head.pt format). The pretrained backbone is resolved from the adapter configuration. A local adapter can still require a backbone download if the backbone is not already cached.

The full question format supports multiple decisions and optional criteria:

result = classifier.classify(
    'Please close my account.',
    questions=[{
        'question': 'What is the customer asking for?',
        'choices': [
            {'choice': 'cancellation', 'criteria': 'Requests to close an account.'},
            {'choice': 'none', 'criteria': 'None of the other choices apply.'},
        ],
    }],
    version='intent',
)

Versions: general, intent, sentiment, triage, cause, codify, complaints, stars. Selecting a new version checks the Hugging Face cache, downloads missing files, then loads the adapter and its scoring head onto the shared backbone. Selection and inference are serialized to keep the adapter and head paired. A failed download raises an error; it never silently falls back to another model.

classifier.load_adapter('custom', checkpoint='./checkpoints/my-adapter')
result = classifier.classify('example', choices=['yes', 'no'], version='custom')

Use model_repo=, revision= and cache_dir= or the THECLASSIFIER_MODEL_REPO environment variable to override download settings. Pin revision to a commit for reproducibility. offline=True disables network loading, including the base model. Pre-download adapters with:

theclassifier download --version general
theclassifier download --version intent

These commands download adapters and tokenizers; initialize the classifier once online to cache the backbone as well. The repository layout is <version>/adapter/, <version>/tokenizer/, <version>/head.safetensors.

Default maximum input length is 2,048 tokens per candidate; longer inputs retain the first token and tail, matching the existing prompt format. Set max_length explicitly for longer records. batch_size controls inference candidate chunks. This initial portable runtime uses direct candidate scoring, not shared-prefix KV caching. It has not been benchmarked against the optimized production server.

Serve on CPU or GPU

theclassifier serve --checkpoint ./checkpoints/general --device cpu
# On a CUDA machine:
API_KEY=your-serving-key theclassifier serve --device cuda --host 0.0.0.0

Default binding is 127.0.0.1:8090. Binding outside localhost requires API_KEY. Use one process per GPU; run behind a TLS proxy for remote clients. This standalone server provides inference and authentication, not the hosted service's billing, quotas or rate limiting. Set resource limits in your deployment.

curl http://127.0.0.1:8090/v1/classify \
  -H 'Content-Type: application/json' \
  -d '{"state":"Close my account","questions":[{"question":"Intent?","choices":[{"choice":"cancel"},{"choice":"none"}]}]}'

Include Authorization: Bearer YOUR_KEY when API_KEY is configured. Responses contain results (with question, selected, detailed_results), version, input_tokens and latency_ms.

Fine-tune with LoRA

Each JSONL line is one complete choice menu, with an exact matching answer label:

{"state":"Close my account","question":"Intent?","choices":[{"choice":"cancel","criteria":"Requests to close an account"},{"choice":"none"}],"answer":"cancel"}
# Continue the general checkpoint to create a specialist:
theclassifier finetune --version general --data train.jsonl \
  --output checkpoints/my-adapter --device cuda

# Train an adapter and scoring head from the base model:
theclassifier finetune --base-model Qwen/Qwen3-1.7B-Base \
  --data train.jsonl --output checkpoints/from-base --rank 16 --alpha 16

Or in Python:

from theclassifier import finetune

report = finetune(
    'train.jsonl', 'checkpoints/my-adapter',
    checkpoint='checkpoints/general', device='auto', epochs=1,
    learning_rate=2e-4, grad_accum=16,
)

Training freezes the backbone, learns LoRA updates to q_proj and v_proj, and trains a float32 scalar head using cross-entropy across each menu. Whole menus are processed together, gradients accumulate between menus, and gradient checkpointing reduces activation memory. Large choice menus may still require smaller input lengths or more memory. CPU training is supported but slow.

New adapters default to rank 16, alpha 16 and dropout 0.05. Continuing an existing adapter preserves its LoRA configuration. A new AdamW optimizer is created with a constant learning rate; this is fine-tuning, not exact training-state resume. The optimizer configuration differs from the original research runs. Output folders must be empty. Saved files include the adapter, head, tokenizer, classifier.json and training.json. Keep evaluation data separate and evaluate before using a new checkpoint; the training loss is not a quality benchmark.

Development and release

cd python
pip install -e '.[serve,dev]'
pytest
python -m build
python -m twine check dist/*
# With PyPI credentials configured locally:
python -m twine upload dist/*

The tests build a tiny random Qwen model locally: no production weights, private data or external inference service is needed. All six tests passed on the NVIDIA workstation, including CUDA inference. Real v1.1 checkpoints were also smoke-tested on CUDA, and the general model was checked on CPU against the original scoring implementation. These checks establish execution compatibility, not accuracy. Publishing the package does not upload model weights. See RELEASE.md for the checkpoint and publication steps.

Metadata

Release files for theclassifier 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for theclassifier 0.1.0
File Size Uploaded
theclassifier-0.1.0.tar.gz 22.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for theclassifier 0.1.0
File Interpreter ABI Platform
theclassifier-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 42.1 kB

Release files / theclassifier-0.1.0.tar.gz

Download URL theclassifier-0.1.0.tar.gz
Size 22.9 kB
Tags Source
SHA-256 checksum
How to use checksums
7d21dda2ec312a78d8bc9f7e1b43dfcf33b9da77479f5eeb8d0e77e1b7b96b77
BLAKE2b-256 checksum
How to use checksums
7c6a2f4a0e08aa61dd9f19a299125d4f557aef98c19651c65c4ef667a9747ea1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release files / theclassifier-0.1.0-py3-none-any.whl

Download URL theclassifier-0.1.0-py3-none-any.whl
Size 19.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
95dd10becbae622f1f690a70f94fe0e1b1751c230469cdbd668aac283d732261
BLAKE2b-256 checksum
How to use checksums
287bfdeece50494d58fa4dd045b51e9f91ffb11f2577e6ea4437e91b06730834
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page