Skip to main content

cureta

Unified Data Curation Framework — enrich text datasets at scale with a composable pipeline of feature-extraction Taggers.


Install

pip install cureta              # core: heuristic taggers, in-process tagging
pip install "cureta[ray]"       # + distributed pipelines (the `cureta run` CLI)
pip install "cureta[ray,llm]"   # + LLM / GPU taggers

Requires: Python ≥ 3.10

Developing on cureta itself? See the contributor setup guide for an editable install from source.


Quick example

Cureta runs a pipeline — a list of taggers that you define. Create one as pipelines/my_pipeline/pipeline_config.yaml:

pipeline_name: "my_pipeline"

tagger_stages:
  - tagger_name: "cureta_id"     # stable per-document content hash
  - tagger_name: "num_words"     # word count
  - tagger_name: "line_stats"    # per-line statistics

Tag a single string in-process with quick_tag:

from cureta import quick_tag

tags = quick_tag(
    "The quick brown fox jumps over the lazy dog.",
    pipeline_config="pipelines/my_pipeline/pipeline_config.yaml",
)
# {'cureta_id': {'cureta_id': 'ef537f2...'}, 'num_words': {'num_words': 10}, ...}

Or run the same pipeline over a whole dataset from the command line:

# Create a small sample dataset (JSONL with a "text" field per line)
python -c "
import json, pathlib
rows = [{'text': f'This is sample document number {i}.'} for i in range(200)]
pathlib.Path('sample.jsonl').write_text('\n'.join(json.dumps(r) for r in rows))
print('Created sample.jsonl with 200 rows')
"

# Resolves pipelines/my_pipeline/pipeline_config.yaml in the current directory
cureta run --pipeline my_pipeline --dataset ./sample.jsonl --limit 100

For HuggingFace datasets see docs/how_to/load_huggingface_data.md


Documentation

Tutorial: Your first pipeline Get a working result in 5 minutes
Tutorial: Writing a custom tagger Build and run your own tagger
Tutorial: Deploying on cloud Run on GPU cluster with SkyPilot
Reference: CLI cureta command reference
Reference: Python API run_pipeline, quick_tag, Pipeline, ...
Reference: Taggers All 45 built-in taggers with output schemas
Concepts: Execution model SIMD vs SDMI, two-phase model, Ray Data
Writing taggers Tier-1 and Tier-2 tagger authoring guide

Examples

01_quickstart Local data, four heuristic taggers, Parquet output
02_custom_tagger Write and run a custom tagger
03_huggingface Load from HuggingFace, use text_column
04_programmatic_api All four Python API surfaces
05_streaming stream_tagged_batches → in-memory / Kafka

Key concepts

  • Document — canonical data envelope with raw_content, document_id, metadata, and tags
  • Tagger — atomic feature-extraction unit; subclass CPUTagger or GPUTagger, implement process_document (Tier 1) or run (Tier 2)
  • Pipeline — orchestrator that loads config, resolves dependencies, and dispatches taggers

45 built-in taggers ship with the library across three categories:

Heuristic taggers (CPU, no model weights) — word count, line/paragraph/word-length statistics, Gopher/C4/FineWeb quality filters, compression ratio, vocabulary diversity, MinHash/SimHash/exact dedup signals, PII regex, HTML density, Markdown structure, content type, URL parsing, domain blocklist, date extraction, and readability metrics.

Lightweight ML taggers — GlotLID language ID (1665 languages), paragraph-level language ID, FastText quality scoring, programming language detection, and Microsoft Presidio PII detection.

GPU taggers — FineWeb educational quality (0–5 score), domain classification (26 domains), multi-label toxicity, dense sentence embeddings, LLM-as-judge scoring, and perplexity.

See Reference: Taggers for the full list with output schemas.


Contributing

See CONTRIBUTING.md and docs/contributing/setup.md.


License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cureta-0.4.0.tar.gz (328.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cureta-0.4.0-py3-none-any.whl (158.9 kB view details)

Uploaded Python 3

File details

Details for the file cureta-0.4.0.tar.gz.

File metadata

  • Download URL: cureta-0.4.0.tar.gz
  • Upload date:
  • Size: 328.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.12

File hashes

Hashes for cureta-0.4.0.tar.gz
Algorithm Hash digest
SHA256 103744aa8b1d5b9b89096f2c5bea1de684efc9b28915895ccd6c5fc09fc1c9ff
MD5 e8f94b6ebb34fbd77eb927b91a13918e
BLAKE2b-256 2cef988e2a91bfd746f815661e558578e7cbfd7dc22cffb19a6b54b52b05e307

See more details on using hashes here.

File details

Details for the file cureta-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: cureta-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 158.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.12

File hashes

Hashes for cureta-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0872541074b1ea9fc167cc145af25551a9127361fd642e8373dd49b4e2ccaba5
MD5 075022dc64b0543fe35d5689dac62c71
BLAKE2b-256 5f622e2b8c70ff90643203ea596dfbf120fe343902fc9b075be777b0f48a4126

See more details on using hashes here.

Release history Release notifications | RSS feed

0.5.0

2 files

This release

0.4.0 This release

2 files

0.3.3

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page