Skip to main content

corpus-preprocess

Github CI

Utility functions to preprocess Phil. legalese in weasel-based flows:

  1. lexcat-proj; and
  2. lexcat-multi

[!IMPORTANT] Requires private corpus-assets folder and sqlite3 db in citelaws-data to be cloned locally.

- corpus-assets: # folder structure
  - concept: # must be two-level nested patterns.json + q.txt
  - artifact: # single folder patterns.json + q.txt
  - text: # each file is a .txt

Language customization

Assuming familiarity with spacy:

nlp.tokenizer = customize_tokenizer(nlp, special_token_rules) # custom tokenizer
ruler = nlp.add_pipe(
    "span_ruler",
    config={
        "spans_key": "ruler",
        "phrase_matcher_attr": "LOWER",
        "spans_filter": {"@misc": "spacy.first_longest_spans_filter.v1"}, # longest spans only
    },
)
ruler.add_patterns(patterns)  # created patterns from this library and corpus-assets

[!NOTE] Loading model with 130k pattern lines takes ~2 min.

Training data

Concept spans

for folder in get_concepts(asset_dir.joinpath("concept")):
    bn = DocBin()
    # use q.txt as queries to the db
    # number of segments per q.txt to fetch
    docs = apply_concept_q_filter(nlp, db_file, filter_path=folder, max_segments=500)
    for doc in docs:
        bn.add(doc)
    bn.to_disk(asset_dir.joinpath(f"train/{folder.stem}.spacy"))

Each concept_dir contains subtopics:

- corpus-assets: # folder structure
  - concept: # must be two-level nested
    - political: # main subject category
        - bill_of_rights: # sub-topic
            - patterns.json # contains matcher files
            - q.txt # contains lines which can be used to query the database

Because of this structure, it's possible to train a textcat_multilabel component:

textcat_options = [concept["id"].split("/")[0] for concept in concept_patterns]

@Language.factory(name="add_cats_from_spans")
class AddTextCatComponent:
    def __init__(self, nlp: Language, name: str, options: list[str]):
        self.nlp = nlp
        self.options = options

    def __call__(self, doc) -> Doc:
        doc.cats = {op: 0.0 for op in self.options}
        for span in doc.spans["sc"]:
            if span.id:  # some spans won't have an id
                value = self.nlp.vocab.strings[span.id]
                if "/" in value:  # e.g. political/bill_of_rights
                    main_topic = value.split("/")[0]  # just political
                    if main_topic in self.options:
                        if doc.cats[main_topic] == 0.0:
                            doc.cats[main_topic] = 1.0
        return doc

Non-concept spans

Although patterns from set_patterns() are included in the constructed nlp object, can ensure that a certain of rows (filter_count) are fetched from the database that have spans which are labeled title and/or serial, etc.

for label in {"unit", "ref", "serial", "title", "axiom", "date", "juridical"}:
    bn = DocBin()
    docs = apply_label_filter(nlp, db_file, filter_labels={label}, filter_count=1500)
    for doc in docs:
        bn.add(doc)
    bn.to_disk(asset_dir.joinpath(f"train/{label}.spacy"))

Release files for corpus-preprocess 0.0.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for corpus-preprocess 0.0.7
File Size Uploaded
corpus_preprocess-0.0.7.tar.gz 21.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for corpus-preprocess 0.0.7
File Interpreter ABI Platform
corpus_preprocess-0.0.7-py3-none-any.whl Python 3 none any Details

Total release size: 49.2 kB

Release files / corpus_preprocess-0.0.7.tar.gz

Download URL corpus_preprocess-0.0.7.tar.gz
Size 21.9 kB
Tags Source
SHA-256 checksum
How to use checksums
41bc54d0f12dfc3fa9ba4ba7b0f0dec86459d7a3f5dbfea07c67c77ead045e6f
BLAKE2b-256 checksum
How to use checksums
eb3f5293b52d7734940cbffc9fee4b41994c4bb5ab58527c7a286bf8e6b9e335
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.7.1 CPython/3.10.6 Darwin/23.2.0

Release files / corpus_preprocess-0.0.7-py3-none-any.whl

Download URL corpus_preprocess-0.0.7-py3-none-any.whl
Size 27.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5d83ebe3e7892ed7759b94877a02e07229641cfcb2a6d4976e7f9e3d136da60d
BLAKE2b-256 checksum
How to use checksums
627988d8f73fe64988d87e5b43972636d6c542c80002bb7c179b9dc773c83979
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.7.1 CPython/3.10.6 Darwin/23.2.0

Release history Release notifications | RSS feed

This release

0.0.7 This release

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page