cleanllm
Streaming JSONL cleaner for LLM fine-tuning datasets. Minimal dependencies, memory-safe, and fast — processes files line-by-line without loading them into memory.
What it does
cleanllm gives you a pipeline for cleaning, validating, and profiling JSONL datasets before fine-tuning:
raw.jsonl → scan → fix → dedup → validate → stats → audit bundle → shards
Every step is streaming (no full-file load), resumable, and produces machine-readable JSON reports for CI gating.
Install
pip install cleanllm
Or from source:
git clone https://github.com/verma8076/cleanllm
cd cleanllm
pip install -e .
For GPU-accelerated scoring, semantic deduplication, and PII redaction:
pip install cleanllm[gpu] # IFD scoring (GPT-2 forward passes)
pip install cleanllm[reward] # reward model scoring (DeBERTa RM)
pip install cleanllm[semantic] # semantic dedup + decontaminate (sentence-transformers)
pip install cleanllm[pii] # spaCy NER for redact-pii
pip install cleanllm[gpu,reward,semantic,pii] # everything
GPU Benchmark
Measured on Google Colab Tesla T4 (15.6 GB VRAM), 500 records, gpt2 (117M params):
| Feature | Device | rec/sec |
|---|---|---|
IFD scoring (ifd-score) |
CPU | 2.9 |
IFD scoring (ifd-score) |
T4 GPU | 43.1 — 14.9× faster |
| IFD scoring via CLI file | CPU | 2.6 |
| IFD scoring via CLI file | T4 GPU | 32.1 — 12.3× faster |
| Semantic dedup | CPU | 25.3 |
| Semantic dedup | T4 GPU | 68.6 — 2.7× faster |
Semantic model: all-MiniLM-L6-v2. Run benchmark_colab.ipynb to reproduce.
Quickstart
# Scan for issues
cleanllm scan data.jsonl
# Fix: remove URLs, normalize whitespace, redact forbidden patterns
cleanllm fix data.jsonl -o data.cleaned.jsonl
# Deduplicate by prompt content
cleanllm dedup data.cleaned.jsonl -o data.dedup.jsonl --by prompt
# Profile the cleaned dataset
cleanllm stats data.dedup.jsonl --report-json stats.json
# Gate in CI: fail if invalid rows increased
cleanllm gate --compare compare.json --rules gate_rules.json
CLI reference
scan
Streaming scan for issues — invalid JSON, missing keys, URLs, forbidden patterns, language distribution, duplicate estimate.
cleanllm scan data.jsonl
cleanllm scan data.jsonl --report-json scan_report.json --dup-estimate
cleanllm scan data.jsonl --preset cp_portable
fix
Remove URLs, normalize whitespace, redact or drop rows with forbidden patterns.
cleanllm fix data.jsonl -o cleaned.jsonl
cleanllm fix data.jsonl -o cleaned.jsonl --drop-on forbidden_pattern --drop-on invalid_json
cleanllm fix data.jsonl -o cleaned.jsonl --preset cpp17_clean --report-json fix_report.json
Drop rules: invalid_json, missing_required_keys, forbidden_pattern, empty_assistant, placeholder, repetitive_response, bad_conversation.
Note on
empty_assistant: By default this drops assistant responses shorter than 20 characters — calibrated for code datasets where very short responses are almost always errors. For text/chat datasets, set--min-assistant-chars 1to only drop truly blank responses.
validate
Schema validation, line by line. Exit code 0 only if all rows pass.
cleanllm validate data.jsonl --schema basic_sft
cleanllm validate data.jsonl --schema cp_sft_v1
| Schema | Required fields |
|---|---|
basic_sft |
id, messages (list of role/content dicts) |
cp_sft_v1 |
id, source, problem_id, messages, tests (non-empty, with input/output) |
dedup
First-occurrence deduplication — by full record, prompt (system+user), or code (assistant).
cleanllm dedup data.jsonl -o deduped.jsonl --by record
cleanllm dedup data.jsonl -o deduped.jsonl --by prompt --normalized
cleanllm dedup data.jsonl -o deduped.jsonl --by code --report-json dedup_report.json
stats
Single-pass profiler: distributions, structural stats, schema counts, response lengths, language distribution.
cleanllm stats data.jsonl
cleanllm stats data.jsonl --schema cp_sft_v1 --keys source,difficulty_bucket --top-k 20
cleanllm stats data.jsonl --report-json stats.json
compare
Diff two stats reports to catch regressions between dataset versions.
cleanllm compare old_stats.json new_stats.json
cleanllm compare old_stats.json new_stats.json --report-json compare.json
cleanllm compare old.jsonl new.jsonl --from-jsonl --schema cp_sft_v1
gate
CI-friendly quality gating. Nonzero exit on failures.
cleanllm gate --stats stats.json --rules gate_rules.json
cleanllm gate --compare compare.json --rules gate_rules.json --strict
cleanllm gate --compare compare.json --inline-rule "counts_diff.invalid_json_rows.delta<=0"
Gate rules JSON:
{
"version": 1,
"mode": "compare",
"rules": [
{"name": "no_new_invalid", "metric": "counts_diff.invalid_json_rows.delta", "op": "<=", "value": 0},
{"name": "enough_valid", "metric": "counts_diff.valid_json_rows.new", "op": ">=", "value": 1000}
]
}
Supported ops: ==, !=, <, <=, >, >=. Severities: error (default), warn.
run
Execute a JSON-defined multi-step pipeline with variable substitution.
cleanllm run --config pipeline.json
cleanllm run --config pipeline.json --set input_path=data.jsonl --set outdir=out/v2
cleanllm run --config pipeline.json --dry-run
Supported step types: fix, validate, dedup, stats, audit, sample, shard, manifest, scan, compare.
sample
Reservoir sampling — random or stratified, deterministic with --seed.
cleanllm sample data.jsonl -o sample.jsonl -n 500 --seed 42
cleanllm sample data.jsonl -o sample.jsonl -n 500 --stratify source,difficulty_bucket
audit
Build a reproducible audit bundle in one command: sampled JSONL + CSV review index (with original line numbers) + summary + manifest.
cleanllm audit data.jsonl --outdir audit_bundle -n 200 --seed 42
cleanllm audit data.jsonl --outdir audit_bundle -n 200 --stratify source --schema cp_sft_v1
Bundle contents: audit_sample.jsonl, audit_index.csv, audit_summary.json, AUDIT_README.md, manifest.json.
shard / manifest
cleanllm shard data.jsonl --outdir shards --size 5000 --gzip
cleanllm manifest shards -o manifest.json
convert
Convert a JSONL file between sharegpt, alpaca, and chatml formats.
cleanllm convert data.jsonl -o converted.jsonl --from sharegpt --to chatml
cleanllm convert data.jsonl -o converted.jsonl --from alpaca --to sharegpt
Supported formats: sharegpt (conversations list), alpaca (instruction/output), chatml (messages list).
merge
Merge multiple JSONL files into one, with optional deduplication.
cleanllm merge a.jsonl b.jsonl c.jsonl -o merged.jsonl
cleanllm merge a.jsonl b.jsonl -o merged.jsonl --dedup
split
Split a JSONL file into train and val sets.
cleanllm split data.jsonl --outdir splits/
cleanllm split data.jsonl --outdir splits/ --ratio 0.95 --seed 42 --no-shuffle
Outputs <basename>_train.jsonl and <basename>_val.jsonl in the output directory. Default ratio is 0.9 (90% train).
ifd-score
Score each record with real Instruction-Following Difficulty (IFD) — PPL(response|instruction) / PPL(response alone). Higher = harder instruction = more valuable for SFT.
# Requires: pip install cleanllm[gpu]
cleanllm ifd-score data.jsonl -o scored.jsonl
cleanllm ifd-score data.jsonl -o scored.jsonl --device cuda --low-threshold 0.3
Stamps each record with _ifd_score. Use cleanllm filter to cut by threshold.
score-rm
Score each record with a reward model. Uses OpenAssistant/reward-model-deberta-v3-large-v2 (~180 MB) by default — works on CPU, faster on GPU.
# Requires: pip install cleanllm[reward]
cleanllm score-rm data.jsonl -o scored.jsonl
cleanllm score-rm data.jsonl -o scored.jsonl --device cuda --low-threshold 0.0
cleanllm score-rm data.jsonl -o scored.jsonl --model OpenAssistant/reward-model-deberta-v3-large-v2
Stamps each record with _rm_score. Combine with cleanllm filter to keep only high-reward examples.
filter
Filter a pre-scored JSONL by any score field. Use after any scoring command.
cleanllm filter scored.jsonl -o filtered.jsonl --min-ifd 0.3
cleanllm filter scored.jsonl -o filtered.jsonl --min-rm 0.0
cleanllm filter scored.jsonl -o filtered.jsonl --min-ifd 0.3 --min-rm 0.0
cleanllm filter scored.jsonl -o filtered.jsonl --require-field _ifd_score --require-field _rm_score
# Verbosity filters (after score-verbosity)
cleanllm filter scored.jsonl -o filtered.jsonl --max-verbosity 5.0 --max-repetition 0.3 --max-filler 0.4
# DEITA / token noise filters (after score-deita / score-token)
cleanllm filter scored.jsonl -o filtered.jsonl --min-deita 3.0 --max-token-noise 0.2
# Multi-field expression
cleanllm filter scored.jsonl -o filtered.jsonl --score-expr "_ifd_score * _rm_score > 0.5"
# Audit trail — writes every dropped record with _rejection_reason, _rejected_by, _rejected_at
cleanllm filter scored.jsonl -o filtered.jsonl --min-ifd 0.3 --rejected-log dropped.jsonl
decontaminate
Remove or flag training examples that overlap with standard evaluation benchmarks using 8-gram fingerprinting. Streaming, no GPU required.
cleanllm decontaminate data.jsonl -o clean.jsonl
cleanllm decontaminate data.jsonl -o clean.jsonl --benchmark mmlu --benchmark gsm8k
cleanllm decontaminate data.jsonl -o clean.jsonl --mode flag # adds _contaminated field instead of dropping
cleanllm decontaminate data.jsonl -o clean.jsonl --n 10 --no-cache
# Custom benchmark:
cleanllm decontaminate data.jsonl -o clean.jsonl --benchmark-file mybench=eval.jsonl
# Add semantic similarity pass to catch paraphrase contamination (requires pip install cleanllm[semantic]):
cleanllm decontaminate data.jsonl -o clean.jsonl --semantic --semantic-threshold 0.85
Supported benchmarks: mmlu, humaneval, gsm8k, arc (all four checked by default). Benchmark n-gram indexes are cached to ~/.cleanllm/benchmarks/ after the first run. Custom benchmarks require a JSONL with a text field per row.
check-turns
Multi-turn conversation quality checks. Flags role violations, length imbalance, empty turns, and structural issues without modifying the file.
cleanllm check-turns data.jsonl
cleanllm check-turns data.jsonl --report-json checks.json
score-token
Token-level noise mask scoring. Writes _token_noise_score per record — the fraction of tokens flagged as high-noise. Compatible with TRL per-token loss weights.
# Requires: pip install cleanllm[gpu]
cleanllm score-token data.jsonl -o scored.jsonl
cleanllm score-token data.jsonl -o scored.jsonl --device cuda
score-deita
DEITA-style complexity × quality scoring via a local LLM (Ollama). Writes _deita_score, _complexity_score, _quality_score.
# Requires: a running Ollama endpoint
cleanllm score-deita data.jsonl -o scored.jsonl
cleanllm score-deita data.jsonl -o scored.jsonl --model llama3.2 --ollama-url http://localhost:11434
score-verbosity
CPU-only verbosity and repetition scoring. No model, no GPU — single streaming pass.
cleanllm score-verbosity data.jsonl -o scored.jsonl
Stamps three fields per record:
| Field | Meaning |
|---|---|
_verbosity_score |
response / instruction word ratio — high values = suspiciously long responses |
_repetition_score |
fraction of 4-grams in the response that appear more than once |
_filler_score |
fraction of sentences that match filler patterns ("I hope this helps", "In conclusion,", etc.) |
Filter downstream: cleanllm filter scored.jsonl -o out.jsonl --max-verbosity 5.0 --max-repetition 0.3
repair
Auto-fix conversation structure. Streaming, no model.
cleanllm repair data.jsonl -o fixed.jsonl
cleanllm repair data.jsonl -o fixed.jsonl --no-strip-sycophancy --no-strip-filler
cleanllm repair data.jsonl -o fixed.jsonl --collapse-consecutive
Operations (all on by default):
- Strip sycophantic prefixes from assistant turns (
"Certainly! ","Great question! ", …) - Strip trailing filler from assistant turns (
"I hope this helps!","Feel free to ask!", …) - Drop trailing user turns so conversations end cleanly on an assistant message
--collapse-consecutive: merge back-to-back same-role messages into one
select
Score-weighted top-K data selection — the final step of the scoring pipeline.
cleanllm select scored.jsonl -o selected.jsonl --top 5000
cleanllm select scored.jsonl -o selected.jsonl --top 20%
cleanllm select scored.jsonl -o selected.jsonl --top 5000 --sort-by _rm_score
cleanllm select scored.jsonl -o selected.jsonl --top 5000 --score-expr "_ifd_score * _rm_score"
cleanllm select scored.jsonl -o selected.jsonl --top 5000 --diversity-weight 0.3
--top Nor--top N%: absolute count or percentage--sort-by FIELD: sort by a single score field (descending)--score-expr EXPR: safe Python arithmetic expression over record fields--diversity-weight 0.0–1.0: blend score ranking with embedding-based diversity (requires[semantic])
mix
Weighted dataset mixing — combine sources with per-file weights.
cleanllm mix source1.jsonl:0.6 source2.jsonl:0.3 source3.jsonl:0.1 -o mixed.jsonl --total 10000
cleanllm mix source1.jsonl source2.jsonl -o mixed.jsonl --strategy round-robin
file.jsonl:weight— weight is a sampling probability (automatically normalized)--total N: target output record count--strategy weighted|round-robin|interleave--seed N: reproducible shuffle
redact-pii
PII detection and redaction. Regex layer always active; spaCy NER is opt-in.
cleanllm redact-pii data.jsonl -o clean.jsonl --mode redact # replace with [EMAIL], [PHONE], etc.
cleanllm redact-pii data.jsonl -o clean.jsonl --mode flag # add _pii_detected field, keep text
cleanllm redact-pii data.jsonl -o clean.jsonl --mode drop # remove rows that contain PII
# spaCy NER for PERSON, ORG, GPE, LOC (requires pip install cleanllm[pii] && python -m spacy download en_core_web_sm)
cleanllm redact-pii data.jsonl -o clean.jsonl --mode redact --use-spacy
Regex-detected types: EMAIL, PHONE, SSN, CREDIT_CARD, IP_ADDRESS, URL. spaCy adds: PERSON, ORG, GPE, LOC.
preflight
Framework-aware pre-flight validation before you start a training run. Catches schema mismatches, bad role sequences, and missing fields that would silently corrupt training.
cleanllm preflight data.jsonl --framework trl
cleanllm preflight data.jsonl --framework axolotl
cleanllm preflight data.jsonl --framework torchtune
cleanllm preflight data.jsonl --framework unsloth
cleanllm preflight data.jsonl --framework trl --report-json preflight.json
Supported frameworks and what they check:
| Framework | Format checked | Key rules |
|---|---|---|
trl |
messages list |
roles: user/assistant/system/tool/ipython; no back-to-back same role; system must be first |
axolotl |
messages OR conversations OR alpaca |
auto-detects format, delegates to appropriate validator |
unsloth |
same as axolotl | — |
torchtune |
messages list |
roles: user/assistant/system/ipython |
Exit code 1 if any rows fail — pipe-safe for CI.
recipes
Bootstrap pipelines and gate rules from built-in templates.
cleanllm recipes list
cleanllm recipes show cp_pipeline_cp_portable
cleanllm recipes write cp_bundle --outdir bootstrap/
Built-in recipes: cp_pipeline_basic, cp_pipeline_cp_portable, cp_pipeline_fast_audit, gate_stats_basic, gate_compare_basic, gate_compare_strict, cp_bundle.
Python API
from cleanllm import (
scan_jsonl, fix_jsonl, FixRules,
dedup_jsonl, validate_jsonl, stats_jsonl,
sample_jsonl, audit_bundle,
shard_jsonl, make_manifest,
download_from_hub, detect_hf_schema,
)
from cleanllm.convert import convert_jsonl
from cleanllm.merge import merge_jsonl
from cleanllm.split import split_jsonl
# Scan
report = scan_jsonl("data.jsonl")
# Fix (code dataset)
rules = FixRules(
drop_on={"forbidden_pattern", "empty_assistant"},
max_tokens=4096,
keep_language="python",
)
summary = fix_jsonl("data.jsonl", "cleaned.jsonl", rules)
# Fix (text/chat dataset — only drop truly blank responses)
rules = FixRules(drop_on={"empty_assistant"}, min_assistant_chars=1, forbidden_patterns=[])
# Dedup
result = dedup_jsonl("cleaned.jsonl", "deduped.jsonl", by="prompt", normalized=True)
# Stats
stats = stats_jsonl("deduped.jsonl", schema="cp_sft_v1", keys=["source", "difficulty_bucket"])
# Sample + audit
sample_jsonl("deduped.jsonl", "sample.jsonl", num_rows=200, seed=42)
audit_bundle("deduped.jsonl", "audit_bundle", num_rows=200, seed=42, stratify=["source"])
# Shard + manifest
shard_jsonl("deduped.jsonl", "shards", shard_size=5000, gzip_output=True)
make_manifest("shards", "manifest.json")
# Convert between formats
convert_jsonl("data.jsonl", "out.jsonl", from_fmt="sharegpt", to_fmt="chatml")
# Merge + split
merge_jsonl(["a.jsonl", "b.jsonl"], "merged.jsonl", dedup=True)
split_jsonl("merged.jsonl", "splits/", ratio=0.9, seed=42)
# Download from HuggingFace Hub (requires pip install cleanllm[hf])
result = download_from_hub("HuggingFaceH4/ultrachat_200k", "data.jsonl", split="train_sft")
# IFD scoring (requires pip install cleanllm[gpu])
from cleanllm.gpu import score_ifd_jsonl
result = score_ifd_jsonl("data.jsonl", "scored.jsonl", device="cuda", low_threshold=0.3)
# Reward model scoring (requires pip install cleanllm[reward])
from cleanllm.gpu import score_rm_jsonl
result = score_rm_jsonl("data.jsonl", "scored.jsonl", device="cuda", low_threshold=0.0)
# Benchmark decontamination
from cleanllm.decontaminate import decontaminate_jsonl
result = decontaminate_jsonl("data.jsonl", "clean.jsonl", benchmarks=["mmlu", "gsm8k"])
# Framework pre-flight validation
from cleanllm.validate import preflight_jsonl
result = preflight_jsonl("data.jsonl", framework="trl")
# Verbosity scoring (CPU-only, no deps)
from cleanllm.verbosity import score_verbosity_jsonl
result = score_verbosity_jsonl("data.jsonl", "scored.jsonl")
# Repair conversation structure
from cleanllm.repair import repair_jsonl
result = repair_jsonl("data.jsonl", "fixed.jsonl", strip_sycophancy=True, strip_filler=True)
# PII detection and redaction
from cleanllm.pii import redact_pii_jsonl
result = redact_pii_jsonl("data.jsonl", "clean.jsonl", mode="redact")
result = redact_pii_jsonl("data.jsonl", "clean.jsonl", mode="flag")
# Score-weighted selection
from cleanllm.select import select_jsonl
result = select_jsonl("scored.jsonl", "selected.jsonl", top=5000, sort_by="_rm_score")
result = select_jsonl("scored.jsonl", "selected.jsonl", top="20%", score_expr="_ifd_score * _rm_score")
# Weighted dataset mixing
from cleanllm.mix import mix_jsonl
result = mix_jsonl(
[("source1.jsonl", 0.6), ("source2.jsonl", 0.4)],
"mixed.jsonl",
total=10000,
strategy="weighted",
seed=42,
)
Presets
| Preset | Description |
|---|---|
general |
URL removal + whitespace normalization, no domain-specific forbidden patterns |
security_scan |
Redacts secrets: AWS keys, GitHub tokens, API keys, private keys |
pii_scan |
Redacts PII: emails, US phone numbers, SSNs, credit cards, IPv4 addresses |
cpp17_clean |
URL removal + whitespace normalization + redact C++ portability issues |
cp_portable |
Strict CP portability — drops rows with forbidden patterns |
deterministic_only |
Drops rows with non-deterministic APIs (rand(), random_device, etc.) |
Defaults
- Required keys:
id,messages - Forbidden patterns (default): none — use
--preset cpp17_cleanor--preset cp_portablefor CP datasets empty_assistantthreshold: 20 characters (responses shorter than this are flagged as empty)
CP datasets: To apply competitive-programming forbidden patterns (
freopen,ifstream,bits/extc++.h, etc.) use a preset:cleanllm fix data.jsonl -o out.jsonl --preset cp_portable. In Python, passforbidden_patterns=list(DEFAULT_FORBIDDEN_PATTERNS)explicitly.
Data format
cleanllm expects JSONL where each line is a JSON object. The default schema (cp_sft_v1) requires:
{
"id": "unique-id",
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]
}
Optional fields: source, difficulty_bucket, problem_id, tests.
Development
pip install -e .[dev]
pytest
python -m build
twine check dist/*
See RELEASE_CHECKLIST.md for the full release workflow.
License
MIT
Release files for cleanllm 4.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cleanllm-4.0.0.tar.gz | 274.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cleanllm-4.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 399.3 kB
Release files / cleanllm-4.0.0.tar.gz
| Download URL | cleanllm-4.0.0.tar.gz |
|---|---|
| Size | 274.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d6755e04707f35342f42d91f544d11ec042d2153b2412ebf89f6d2967f8b84d0
|
|
BLAKE2b-256 checksum How to use checksums |
a1f9d0d4c34f63133fd83d55764460f5b2f1d5224b3a1e4000d531dbb3a9da4d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.13.7
|
Release files / cleanllm-4.0.0-py3-none-any.whl
| Download URL | cleanllm-4.0.0-py3-none-any.whl |
|---|---|
| Size | 125.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1a3afce60776bdcee98d6db5f48e8a1c6383b0d974fe183faa80afa80603fc58
|
|
BLAKE2b-256 checksum How to use checksums |
8ec7242f12448da61482ce82faf5e6d682830b4ed634696b5cb9a412994b9de5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.13.7
|