Skip to main content

cleanllm

Streaming JSONL cleaner for LLM fine-tuning datasets. Minimal dependencies, memory-safe, and fast — processes files line-by-line without loading them into memory.

PyPI Python License: MIT


What it does

cleanllm gives you a pipeline for cleaning, validating, and profiling JSONL datasets before fine-tuning:

raw.jsonl → scan → fix → dedup → validate → stats → audit bundle → shards

Every step is streaming (no full-file load), resumable, and produces machine-readable JSON reports for CI gating.


Install

pip install cleanllm

Or from source:

git clone https://github.com/verma8076/cleanllm
cd cleanllm
pip install -e .

For GPU-accelerated scoring, semantic deduplication, and PII redaction:

pip install cleanllm[gpu]              # IFD scoring (GPT-2 forward passes)
pip install cleanllm[reward]           # reward model scoring (DeBERTa RM)
pip install cleanllm[semantic]         # semantic dedup + decontaminate (sentence-transformers)
pip install cleanllm[pii]              # spaCy NER for redact-pii
pip install cleanllm[gpu,reward,semantic,pii]  # everything

GPU Benchmark

Measured on Google Colab Tesla T4 (15.6 GB VRAM), 500 records, gpt2 (117M params):

Feature Device rec/sec
IFD scoring (ifd-score) CPU 2.9
IFD scoring (ifd-score) T4 GPU 43.1 — 14.9× faster
IFD scoring via CLI file CPU 2.6
IFD scoring via CLI file T4 GPU 32.1 — 12.3× faster
Semantic dedup CPU 25.3
Semantic dedup T4 GPU 68.6 — 2.7× faster

Semantic model: all-MiniLM-L6-v2. Run benchmark_colab.ipynb to reproduce.


Quickstart

# Scan for issues
cleanllm scan data.jsonl

# Fix: remove URLs, normalize whitespace, redact forbidden patterns
cleanllm fix data.jsonl -o data.cleaned.jsonl

# Deduplicate by prompt content
cleanllm dedup data.cleaned.jsonl -o data.dedup.jsonl --by prompt

# Profile the cleaned dataset
cleanllm stats data.dedup.jsonl --report-json stats.json

# Gate in CI: fail if invalid rows increased
cleanllm gate --compare compare.json --rules gate_rules.json

CLI reference

scan

Streaming scan for issues — invalid JSON, missing keys, URLs, forbidden patterns, language distribution, duplicate estimate.

cleanllm scan data.jsonl
cleanllm scan data.jsonl --report-json scan_report.json --dup-estimate
cleanllm scan data.jsonl --preset cp_portable

fix

Remove URLs, normalize whitespace, redact or drop rows with forbidden patterns.

cleanllm fix data.jsonl -o cleaned.jsonl
cleanllm fix data.jsonl -o cleaned.jsonl --drop-on forbidden_pattern --drop-on invalid_json
cleanllm fix data.jsonl -o cleaned.jsonl --preset cpp17_clean --report-json fix_report.json

Drop rules: invalid_json, missing_required_keys, forbidden_pattern, empty_assistant, placeholder, repetitive_response, bad_conversation.

Note on empty_assistant: By default this drops assistant responses shorter than 20 characters — calibrated for code datasets where very short responses are almost always errors. For text/chat datasets, set --min-assistant-chars 1 to only drop truly blank responses.

validate

Schema validation, line by line. Exit code 0 only if all rows pass.

cleanllm validate data.jsonl --schema basic_sft
cleanllm validate data.jsonl --schema cp_sft_v1
Schema Required fields
basic_sft id, messages (list of role/content dicts)
cp_sft_v1 id, source, problem_id, messages, tests (non-empty, with input/output)

dedup

First-occurrence deduplication — by full record, prompt (system+user), or code (assistant).

cleanllm dedup data.jsonl -o deduped.jsonl --by record
cleanllm dedup data.jsonl -o deduped.jsonl --by prompt --normalized
cleanllm dedup data.jsonl -o deduped.jsonl --by code --report-json dedup_report.json

stats

Single-pass profiler: distributions, structural stats, schema counts, response lengths, language distribution.

cleanllm stats data.jsonl
cleanllm stats data.jsonl --schema cp_sft_v1 --keys source,difficulty_bucket --top-k 20
cleanllm stats data.jsonl --report-json stats.json

compare

Diff two stats reports to catch regressions between dataset versions.

cleanllm compare old_stats.json new_stats.json
cleanllm compare old_stats.json new_stats.json --report-json compare.json
cleanllm compare old.jsonl new.jsonl --from-jsonl --schema cp_sft_v1

gate

CI-friendly quality gating. Nonzero exit on failures.

cleanllm gate --stats stats.json --rules gate_rules.json
cleanllm gate --compare compare.json --rules gate_rules.json --strict
cleanllm gate --compare compare.json --inline-rule "counts_diff.invalid_json_rows.delta<=0"

Gate rules JSON:

{
  "version": 1,
  "mode": "compare",
  "rules": [
    {"name": "no_new_invalid", "metric": "counts_diff.invalid_json_rows.delta", "op": "<=", "value": 0},
    {"name": "enough_valid",   "metric": "counts_diff.valid_json_rows.new",    "op": ">=", "value": 1000}
  ]
}

Supported ops: ==, !=, <, <=, >, >=. Severities: error (default), warn.

run

Execute a JSON-defined multi-step pipeline with variable substitution.

cleanllm run --config pipeline.json
cleanllm run --config pipeline.json --set input_path=data.jsonl --set outdir=out/v2
cleanllm run --config pipeline.json --dry-run

Supported step types: fix, validate, dedup, stats, audit, sample, shard, manifest, scan, compare.

sample

Reservoir sampling — random or stratified, deterministic with --seed.

cleanllm sample data.jsonl -o sample.jsonl -n 500 --seed 42
cleanllm sample data.jsonl -o sample.jsonl -n 500 --stratify source,difficulty_bucket

audit

Build a reproducible audit bundle in one command: sampled JSONL + CSV review index (with original line numbers) + summary + manifest.

cleanllm audit data.jsonl --outdir audit_bundle -n 200 --seed 42
cleanllm audit data.jsonl --outdir audit_bundle -n 200 --stratify source --schema cp_sft_v1

Bundle contents: audit_sample.jsonl, audit_index.csv, audit_summary.json, AUDIT_README.md, manifest.json.

shard / manifest

cleanllm shard data.jsonl --outdir shards --size 5000 --gzip
cleanllm manifest shards -o manifest.json

convert

Convert a JSONL file between sharegpt, alpaca, and chatml formats.

cleanllm convert data.jsonl -o converted.jsonl --from sharegpt --to chatml
cleanllm convert data.jsonl -o converted.jsonl --from alpaca --to sharegpt

Supported formats: sharegpt (conversations list), alpaca (instruction/output), chatml (messages list).

merge

Merge multiple JSONL files into one, with optional deduplication.

cleanllm merge a.jsonl b.jsonl c.jsonl -o merged.jsonl
cleanllm merge a.jsonl b.jsonl -o merged.jsonl --dedup

split

Split a JSONL file into train and val sets.

cleanllm split data.jsonl --outdir splits/
cleanllm split data.jsonl --outdir splits/ --ratio 0.95 --seed 42 --no-shuffle

Outputs <basename>_train.jsonl and <basename>_val.jsonl in the output directory. Default ratio is 0.9 (90% train).

ifd-score

Score each record with real Instruction-Following Difficulty (IFD) — PPL(response|instruction) / PPL(response alone). Higher = harder instruction = more valuable for SFT.

# Requires: pip install cleanllm[gpu]
cleanllm ifd-score data.jsonl -o scored.jsonl
cleanllm ifd-score data.jsonl -o scored.jsonl --device cuda --low-threshold 0.3

Stamps each record with _ifd_score. Use cleanllm filter to cut by threshold.

score-rm

Score each record with a reward model. Uses OpenAssistant/reward-model-deberta-v3-large-v2 (~180 MB) by default — works on CPU, faster on GPU.

# Requires: pip install cleanllm[reward]
cleanllm score-rm data.jsonl -o scored.jsonl
cleanllm score-rm data.jsonl -o scored.jsonl --device cuda --low-threshold 0.0
cleanllm score-rm data.jsonl -o scored.jsonl --model OpenAssistant/reward-model-deberta-v3-large-v2

Stamps each record with _rm_score. Combine with cleanllm filter to keep only high-reward examples.

filter

Filter a pre-scored JSONL by any score field. Use after any scoring command.

cleanllm filter scored.jsonl -o filtered.jsonl --min-ifd 0.3
cleanllm filter scored.jsonl -o filtered.jsonl --min-rm 0.0
cleanllm filter scored.jsonl -o filtered.jsonl --min-ifd 0.3 --min-rm 0.0
cleanllm filter scored.jsonl -o filtered.jsonl --require-field _ifd_score --require-field _rm_score

# Verbosity filters (after score-verbosity)
cleanllm filter scored.jsonl -o filtered.jsonl --max-verbosity 5.0 --max-repetition 0.3 --max-filler 0.4

# DEITA / token noise filters (after score-deita / score-token)
cleanllm filter scored.jsonl -o filtered.jsonl --min-deita 3.0 --max-token-noise 0.2

# Multi-field expression
cleanllm filter scored.jsonl -o filtered.jsonl --score-expr "_ifd_score * _rm_score > 0.5"

# Audit trail — writes every dropped record with _rejection_reason, _rejected_by, _rejected_at
cleanllm filter scored.jsonl -o filtered.jsonl --min-ifd 0.3 --rejected-log dropped.jsonl

decontaminate

Remove or flag training examples that overlap with standard evaluation benchmarks using 8-gram fingerprinting. Streaming, no GPU required.

cleanllm decontaminate data.jsonl -o clean.jsonl
cleanllm decontaminate data.jsonl -o clean.jsonl --benchmark mmlu --benchmark gsm8k
cleanllm decontaminate data.jsonl -o clean.jsonl --mode flag   # adds _contaminated field instead of dropping
cleanllm decontaminate data.jsonl -o clean.jsonl --n 10 --no-cache
# Custom benchmark:
cleanllm decontaminate data.jsonl -o clean.jsonl --benchmark-file mybench=eval.jsonl
# Add semantic similarity pass to catch paraphrase contamination (requires pip install cleanllm[semantic]):
cleanllm decontaminate data.jsonl -o clean.jsonl --semantic --semantic-threshold 0.85

Supported benchmarks: mmlu, humaneval, gsm8k, arc (all four checked by default). Benchmark n-gram indexes are cached to ~/.cleanllm/benchmarks/ after the first run. Custom benchmarks require a JSONL with a text field per row.

check-turns

Multi-turn conversation quality checks. Flags role violations, length imbalance, empty turns, and structural issues without modifying the file.

cleanllm check-turns data.jsonl
cleanllm check-turns data.jsonl --report-json checks.json

score-token

Token-level noise mask scoring. Writes _token_noise_score per record — the fraction of tokens flagged as high-noise. Compatible with TRL per-token loss weights.

# Requires: pip install cleanllm[gpu]
cleanllm score-token data.jsonl -o scored.jsonl
cleanllm score-token data.jsonl -o scored.jsonl --device cuda

score-deita

DEITA-style complexity × quality scoring via a local LLM (Ollama). Writes _deita_score, _complexity_score, _quality_score.

# Requires: a running Ollama endpoint
cleanllm score-deita data.jsonl -o scored.jsonl
cleanllm score-deita data.jsonl -o scored.jsonl --model llama3.2 --ollama-url http://localhost:11434

score-verbosity

CPU-only verbosity and repetition scoring. No model, no GPU — single streaming pass.

cleanllm score-verbosity data.jsonl -o scored.jsonl

Stamps three fields per record:

Field Meaning
_verbosity_score response / instruction word ratio — high values = suspiciously long responses
_repetition_score fraction of 4-grams in the response that appear more than once
_filler_score fraction of sentences that match filler patterns ("I hope this helps", "In conclusion,", etc.)

Filter downstream: cleanllm filter scored.jsonl -o out.jsonl --max-verbosity 5.0 --max-repetition 0.3

repair

Auto-fix conversation structure. Streaming, no model.

cleanllm repair data.jsonl -o fixed.jsonl
cleanllm repair data.jsonl -o fixed.jsonl --no-strip-sycophancy --no-strip-filler
cleanllm repair data.jsonl -o fixed.jsonl --collapse-consecutive

Operations (all on by default):

  • Strip sycophantic prefixes from assistant turns ("Certainly! ", "Great question! ", …)
  • Strip trailing filler from assistant turns ("I hope this helps!", "Feel free to ask!", …)
  • Drop trailing user turns so conversations end cleanly on an assistant message
  • --collapse-consecutive: merge back-to-back same-role messages into one

select

Score-weighted top-K data selection — the final step of the scoring pipeline.

cleanllm select scored.jsonl -o selected.jsonl --top 5000
cleanllm select scored.jsonl -o selected.jsonl --top 20%
cleanllm select scored.jsonl -o selected.jsonl --top 5000 --sort-by _rm_score
cleanllm select scored.jsonl -o selected.jsonl --top 5000 --score-expr "_ifd_score * _rm_score"
cleanllm select scored.jsonl -o selected.jsonl --top 5000 --diversity-weight 0.3
  • --top N or --top N%: absolute count or percentage
  • --sort-by FIELD: sort by a single score field (descending)
  • --score-expr EXPR: safe Python arithmetic expression over record fields
  • --diversity-weight 0.0–1.0: blend score ranking with embedding-based diversity (requires [semantic])

mix

Weighted dataset mixing — combine sources with per-file weights.

cleanllm mix source1.jsonl:0.6 source2.jsonl:0.3 source3.jsonl:0.1 -o mixed.jsonl --total 10000
cleanllm mix source1.jsonl source2.jsonl -o mixed.jsonl --strategy round-robin
  • file.jsonl:weight — weight is a sampling probability (automatically normalized)
  • --total N: target output record count
  • --strategy weighted|round-robin|interleave
  • --seed N: reproducible shuffle

redact-pii

PII detection and redaction. Regex layer always active; spaCy NER is opt-in.

cleanllm redact-pii data.jsonl -o clean.jsonl --mode redact   # replace with [EMAIL], [PHONE], etc.
cleanllm redact-pii data.jsonl -o clean.jsonl --mode flag     # add _pii_detected field, keep text
cleanllm redact-pii data.jsonl -o clean.jsonl --mode drop     # remove rows that contain PII

# spaCy NER for PERSON, ORG, GPE, LOC (requires pip install cleanllm[pii] && python -m spacy download en_core_web_sm)
cleanllm redact-pii data.jsonl -o clean.jsonl --mode redact --use-spacy

Regex-detected types: EMAIL, PHONE, SSN, CREDIT_CARD, IP_ADDRESS, URL. spaCy adds: PERSON, ORG, GPE, LOC.

preflight

Framework-aware pre-flight validation before you start a training run. Catches schema mismatches, bad role sequences, and missing fields that would silently corrupt training.

cleanllm preflight data.jsonl --framework trl
cleanllm preflight data.jsonl --framework axolotl
cleanllm preflight data.jsonl --framework torchtune
cleanllm preflight data.jsonl --framework unsloth
cleanllm preflight data.jsonl --framework trl --report-json preflight.json

Supported frameworks and what they check:

Framework Format checked Key rules
trl messages list roles: user/assistant/system/tool/ipython; no back-to-back same role; system must be first
axolotl messages OR conversations OR alpaca auto-detects format, delegates to appropriate validator
unsloth same as axolotl
torchtune messages list roles: user/assistant/system/ipython

Exit code 1 if any rows fail — pipe-safe for CI.

recipes

Bootstrap pipelines and gate rules from built-in templates.

cleanllm recipes list
cleanllm recipes show cp_pipeline_cp_portable
cleanllm recipes write cp_bundle --outdir bootstrap/

Built-in recipes: cp_pipeline_basic, cp_pipeline_cp_portable, cp_pipeline_fast_audit, gate_stats_basic, gate_compare_basic, gate_compare_strict, cp_bundle.


Python API

from cleanllm import (
    scan_jsonl, fix_jsonl, FixRules,
    dedup_jsonl, validate_jsonl, stats_jsonl,
    sample_jsonl, audit_bundle,
    shard_jsonl, make_manifest,
    download_from_hub, detect_hf_schema,
)
from cleanllm.convert import convert_jsonl
from cleanllm.merge import merge_jsonl
from cleanllm.split import split_jsonl

# Scan
report = scan_jsonl("data.jsonl")

# Fix (code dataset)
rules = FixRules(
    drop_on={"forbidden_pattern", "empty_assistant"},
    max_tokens=4096,
    keep_language="python",
)
summary = fix_jsonl("data.jsonl", "cleaned.jsonl", rules)

# Fix (text/chat dataset — only drop truly blank responses)
rules = FixRules(drop_on={"empty_assistant"}, min_assistant_chars=1, forbidden_patterns=[])

# Dedup
result = dedup_jsonl("cleaned.jsonl", "deduped.jsonl", by="prompt", normalized=True)

# Stats
stats = stats_jsonl("deduped.jsonl", schema="cp_sft_v1", keys=["source", "difficulty_bucket"])

# Sample + audit
sample_jsonl("deduped.jsonl", "sample.jsonl", num_rows=200, seed=42)
audit_bundle("deduped.jsonl", "audit_bundle", num_rows=200, seed=42, stratify=["source"])

# Shard + manifest
shard_jsonl("deduped.jsonl", "shards", shard_size=5000, gzip_output=True)
make_manifest("shards", "manifest.json")

# Convert between formats
convert_jsonl("data.jsonl", "out.jsonl", from_fmt="sharegpt", to_fmt="chatml")

# Merge + split
merge_jsonl(["a.jsonl", "b.jsonl"], "merged.jsonl", dedup=True)
split_jsonl("merged.jsonl", "splits/", ratio=0.9, seed=42)

# Download from HuggingFace Hub (requires pip install cleanllm[hf])
result = download_from_hub("HuggingFaceH4/ultrachat_200k", "data.jsonl", split="train_sft")

# IFD scoring (requires pip install cleanllm[gpu])
from cleanllm.gpu import score_ifd_jsonl
result = score_ifd_jsonl("data.jsonl", "scored.jsonl", device="cuda", low_threshold=0.3)

# Reward model scoring (requires pip install cleanllm[reward])
from cleanllm.gpu import score_rm_jsonl
result = score_rm_jsonl("data.jsonl", "scored.jsonl", device="cuda", low_threshold=0.0)

# Benchmark decontamination
from cleanllm.decontaminate import decontaminate_jsonl
result = decontaminate_jsonl("data.jsonl", "clean.jsonl", benchmarks=["mmlu", "gsm8k"])

# Framework pre-flight validation
from cleanllm.validate import preflight_jsonl
result = preflight_jsonl("data.jsonl", framework="trl")

# Verbosity scoring (CPU-only, no deps)
from cleanllm.verbosity import score_verbosity_jsonl
result = score_verbosity_jsonl("data.jsonl", "scored.jsonl")

# Repair conversation structure
from cleanllm.repair import repair_jsonl
result = repair_jsonl("data.jsonl", "fixed.jsonl", strip_sycophancy=True, strip_filler=True)

# PII detection and redaction
from cleanllm.pii import redact_pii_jsonl
result = redact_pii_jsonl("data.jsonl", "clean.jsonl", mode="redact")
result = redact_pii_jsonl("data.jsonl", "clean.jsonl", mode="flag")

# Score-weighted selection
from cleanllm.select import select_jsonl
result = select_jsonl("scored.jsonl", "selected.jsonl", top=5000, sort_by="_rm_score")
result = select_jsonl("scored.jsonl", "selected.jsonl", top="20%", score_expr="_ifd_score * _rm_score")

# Weighted dataset mixing
from cleanllm.mix import mix_jsonl
result = mix_jsonl(
    [("source1.jsonl", 0.6), ("source2.jsonl", 0.4)],
    "mixed.jsonl",
    total=10000,
    strategy="weighted",
    seed=42,
)

Presets

Preset Description
general URL removal + whitespace normalization, no domain-specific forbidden patterns
security_scan Redacts secrets: AWS keys, GitHub tokens, API keys, private keys
pii_scan Redacts PII: emails, US phone numbers, SSNs, credit cards, IPv4 addresses
cpp17_clean URL removal + whitespace normalization + redact C++ portability issues
cp_portable Strict CP portability — drops rows with forbidden patterns
deterministic_only Drops rows with non-deterministic APIs (rand(), random_device, etc.)

Defaults

  • Required keys: id, messages
  • Forbidden patterns (default): none — use --preset cpp17_clean or --preset cp_portable for CP datasets
  • empty_assistant threshold: 20 characters (responses shorter than this are flagged as empty)

CP datasets: To apply competitive-programming forbidden patterns (freopen, ifstream, bits/extc++.h, etc.) use a preset: cleanllm fix data.jsonl -o out.jsonl --preset cp_portable. In Python, pass forbidden_patterns=list(DEFAULT_FORBIDDEN_PATTERNS) explicitly.


Data format

cleanllm expects JSONL where each line is a JSON object. The default schema (cp_sft_v1) requires:

{
  "id": "unique-id",
  "messages": [
    {"role": "system",    "content": "..."},
    {"role": "user",      "content": "..."},
    {"role": "assistant", "content": "..."}
  ]
}

Optional fields: source, difficulty_bucket, problem_id, tests.


Development

pip install -e .[dev]
pytest
python -m build
twine check dist/*

See RELEASE_CHECKLIST.md for the full release workflow.


License

MIT

Release files for cleanllm 4.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cleanllm 4.0.0
File Size Uploaded
cleanllm-4.0.0.tar.gz 274.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cleanllm 4.0.0
File Interpreter ABI Platform
cleanllm-4.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 399.3 kB

Release files / cleanllm-4.0.0.tar.gz

Download URL cleanllm-4.0.0.tar.gz
Size 274.3 kB
Tags Source
SHA-256 checksum
How to use checksums
d6755e04707f35342f42d91f544d11ec042d2153b2412ebf89f6d2967f8b84d0
BLAKE2b-256 checksum
How to use checksums
a1f9d0d4c34f63133fd83d55764460f5b2f1d5224b3a1e4000d531dbb3a9da4d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.13.7

Release files / cleanllm-4.0.0-py3-none-any.whl

Download URL cleanllm-4.0.0-py3-none-any.whl
Size 125.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1a3afce60776bdcee98d6db5f48e8a1c6383b0d974fe183faa80afa80603fc58
BLAKE2b-256 checksum
How to use checksums
8ec7242f12448da61482ce82faf5e6d682830b4ed634696b5cb9a412994b9de5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.13.7

Release history Release notifications | RSS feed

This release

4.0.0 This release

2 release files

3.0.0

2 release files

2.0.0

2 release files

1.0.0

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page