Skip to main content

cureta

Unified Data Curation Framework — enrich text datasets at scale with a composable pipeline of feature-extraction Taggers.


Install

pip install cureta              # core: heuristic taggers, in-process tagging
pip install "cureta[ray]"       # + distributed pipelines (the `cureta run` CLI)
pip install "cureta[ray,llm]"   # + LLM / GPU taggers

Requires: Python ≥ 3.10

Developing on cureta itself? See the contributor setup guide for an editable install from source.


Quick example

Cureta runs a pipeline — a list of taggers that you define. Create one as pipelines/my_pipeline/pipeline_config.yaml:

pipeline_name: "my_pipeline"

tagger_stages:
  - tagger_name: "cureta_id"     # stable per-document content hash
  - tagger_name: "num_words"     # word count
  - tagger_name: "line_stats"    # per-line statistics

Tag a single string in-process with quick_tag:

from cureta import quick_tag

tags = quick_tag(
    "The quick brown fox jumps over the lazy dog.",
    pipeline_config="pipelines/my_pipeline/pipeline_config.yaml",
)
# {'cureta_id': {'cureta_id': 'ef537f2...'}, 'num_words': {'num_words': 10}, ...}

Or run the same pipeline over a whole dataset from the command line:

# Create a small sample dataset (JSONL with a "text" field per line)
python -c "
import json, pathlib
rows = [{'text': f'This is sample document number {i}.'} for i in range(200)]
pathlib.Path('sample.jsonl').write_text('\n'.join(json.dumps(r) for r in rows))
print('Created sample.jsonl with 200 rows')
"

# Resolves pipelines/my_pipeline/pipeline_config.yaml in the current directory
cureta run --pipeline my_pipeline --dataset ./sample.jsonl --limit 100

For HuggingFace datasets see docs/how_to/load_huggingface_data.md


Documentation

Tutorial: Your first pipeline Get a working result in 5 minutes
Tutorial: Writing a custom tagger Build and run your own tagger
Tutorial: Deploying on cloud Run on GPU cluster with SkyPilot
Reference: CLI cureta command reference
Reference: Python API run_pipeline, quick_tag, Pipeline, ...
Reference: Taggers All 45 built-in taggers with output schemas
Concepts: Execution model SIMD vs SDMI, two-phase model, Ray Data
Writing taggers Tier-1 and Tier-2 tagger authoring guide

Examples

01_quickstart Local data, four heuristic taggers, Parquet output
02_custom_tagger Write and run a custom tagger
03_huggingface Load from HuggingFace, use text_column
04_programmatic_api All four Python API surfaces
05_streaming stream_tagged_batches → in-memory / Kafka

Key concepts

  • Document — canonical data envelope with raw_content, document_id, metadata, and tags
  • Tagger — atomic feature-extraction unit; subclass CPUTagger or GPUTagger, implement process_document (Tier 1) or run (Tier 2)
  • Pipeline — orchestrator that loads config, resolves dependencies, and dispatches taggers

45 built-in taggers ship with the library across three categories:

Heuristic taggers (CPU, no model weights) — word count, line/paragraph/word-length statistics, Gopher/C4/FineWeb quality filters, compression ratio, vocabulary diversity, MinHash/SimHash/exact dedup signals, PII regex, HTML density, Markdown structure, content type, URL parsing, domain blocklist, date extraction, and readability metrics.

Lightweight ML taggers — GlotLID language ID (1665 languages), paragraph-level language ID, FastText quality scoring, programming language detection, and Microsoft Presidio PII detection.

GPU taggers — FineWeb educational quality (0–5 score), domain classification (26 domains), multi-label toxicity, dense sentence embeddings, LLM-as-judge scoring, and perplexity.

See Reference: Taggers for the full list with output schemas.


Contributing

See CONTRIBUTING.md and docs/contributing/setup.md.


License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cureta-0.3.2.tar.gz (319.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cureta-0.3.2-py3-none-any.whl (156.3 kB view details)

Uploaded Python 3

File details

Details for the file cureta-0.3.2.tar.gz.

File metadata

  • Download URL: cureta-0.3.2.tar.gz
  • Upload date:
  • Size: 319.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for cureta-0.3.2.tar.gz
Algorithm Hash digest
SHA256 26b2daaf27119d7c499a45d2780b7f19bd0068326961c967cb1c652125980d47
MD5 857f7fa9a0bd7c474f243ba872738933
BLAKE2b-256 ded31f04ff7ec1e9bcfa76986efd464af609b16b35f7491b89511c302ddd0c0f

See more details on using hashes here.

File details

Details for the file cureta-0.3.2-py3-none-any.whl.

File metadata

  • Download URL: cureta-0.3.2-py3-none-any.whl
  • Upload date:
  • Size: 156.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for cureta-0.3.2-py3-none-any.whl
Algorithm Hash digest
SHA256 199d71f5f4cb49cae16c897722a702102b7a11d241944de5b88204d5468c9f3b
MD5 4c8b555f3bc067bfde3ebfecd21123c1
BLAKE2b-256 b6be770a6701b7fc1b9477e55e458c69e0cb345d555e0e90f78f3accbcde493c

See more details on using hashes here.

Release history Release notifications | RSS feed

0.5.0

2 files

0.4.0

2 files

0.3.3

2 files

This release

0.3.2 This release

2 files

0.3.1

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page