Skip to main content

DSTPR: Domain-specific Text Pre-processing and Ranking

PyPI version

A high-throughput text preprocessing pipeline designed to filter, segment, and rank core sentences from noisy, plaintext documents.

This package strips away boilerplate, disclaimers, and application instructions before moving this data onto to heavier processing pipelines.


Pipeline Architecture

Instead of running heavy transformer models over every sentence in a document, DSTPR pipes documents through progressively stricter layers:

  1. Heuristic Cleansing & Tokenization: Uses PySBD (Python Sentence Boundary Disambiguation) paired with regular expressions to fix punctuation caused by flattening, then filters out sentences containing stop phrases.
  2. Semantic Section Routing: Utilizes a lightweight encoder (Default: all-MiniLM-L6-v2) to dynamically find structural transitions (i.e., headers) in input texts.
  3. Hybrid Ranking: Scores remaining sentences using a combination of Semantics (cosine similarity to anchors) and Lexical Syntax (regex- and POS-supported detection).
  4. Parallel Execution Engine: Wraps the entire pipeline inside efficient datasets to allow for batch processing at scale, with CPU or GPU.

Installation

To install the package, simply run

pip install dstpr

Usage

High-Throughput Batch Processing

This is the recommended approach for large-scale data pipelines.

from dstpr import ParallelPreprocessingPipeline, JOB_POSTING_PROFILE

pipeline = ParallelPreprocessingPipeline(profile=JOB_POSTING_PROFILE, batch_size=64)

raw_documents = [
    "DOC 1",
    "DOC 2",
    "..."
]

cleaned_documents = pipeline.process(raw_documents, num_workers=8, threshold=0.25)

Advanced Usage

If you want to integrate specific pipeline layers directly into an existing workflow, or tweak the internal parameters, you can import individual modules manually:

from dstpr.cleaners import clean_and_split_chunks
from dstpr.segmenters import SemanticSectionRouter
from dstpr.rankers import HybridTaskRanker
from sentence_transformers import Transformer, SentenceTransformer

embedding_model = SentenceTransformer("all-MiniLM-L6-v2")

text = "Job posting text goes here!"

# Clean and segment sentences
sentences = clean_and_split_chunks(text)

# Route by context
router = SemanticSectionRouter(model=embedding_model)
buckets = router.route_sentences(sentences)
target_sentences = buckets['CORE'] + buckets['REQUIREMENTS']

# Grade and sort text features
ranker = HybridRanker(model=shared_model)
final = ranker.rank_and_filter(target_sentences, bi_cutoff_pct=0.25, final_threshold=0.40)

Parameter Tuning

To adjust the trade-off between strict filtering and execution speed, consider tweaking these variables in ParallelPreprocessingPipeline:

  • threshold (Default: 0.4): Controls how aggressively sentences are discarded. Raising this value toward 0.6, for example, ensures only stronger matching sentences are passed through. Lowering it toward 0.3, on the other hand, acts as a wider net.
  • batch_size (Default: 256): This depends on your hardware! Adjust for best performance.

Creating Custom Domain Profiles (via the provided wizard)

A DomainProfile is required to use DSTPR (see the usage example). Out of the box, dstpr ships with pre-configured configurations for job postings (JOB_POSTINGS_PROFILE). However, you can easily generate an domain-specific pipeline for any specific domain using our built-in interactive configuration wizard.

To spin up the profile creation walkthrough, simply open your terminal and run:

task-profile-wizard

All profiles generated via the terminal wizard are automatically validated and written out as JSON to a local cache: ~/.config/dstpr/profiles/

Metadata

Release files for dstpr 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dstpr 0.2.1
File Size Uploaded
dstpr-0.2.1.tar.gz 16.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dstpr 0.2.1
File Interpreter ABI Platform
dstpr-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 34.2 kB

Release files / dstpr-0.2.1.tar.gz

Download URL dstpr-0.2.1.tar.gz
Size 16.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f024add692bee99251f2c827faf62422f031f15329b2514596f37e661a332efe
BLAKE2b-256 checksum
How to use checksums
f6f01da5548feccb97ea32df8983e8f6103c3401284720f8365670159dbe4bcb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.20

Release files / dstpr-0.2.1-py3-none-any.whl

Download URL dstpr-0.2.1-py3-none-any.whl
Size 17.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4756bda902cead126bb9c9e573142eb7dd73cc0fb8e5206fcf0260ef4d5193e5
BLAKE2b-256 checksum
How to use checksums
cfa54ecb9cd654cafe8eb915f17eb66e642712f127af93368b73d02fbf8bbc6a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.20

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page