breakingnews
Purpose. A broadcast-news transcript arrives as an undifferentiated block of 2,000–17,000 words covering several unrelated stories. This package finds the word offsets where one story ends and the next begins, turns them into one row per story, and lets you put the rows back together again. It exists because content analysis needs a comparable unit: a television transcript is not one article, and treating it as one — or splitting it on speaker turns — gives the wrong denominator. A "boundary" means the broadcast moves to a genuinely different story: new topic, new event, different actors, and explicitly not a change of speaker, correspondent, location or sub-angle within a continuing story. Boundaries only — nothing here labels, classifies or summarises the segments it produces.
Install
pip install breakingnews # scoring, segments, merge, reconcile
pip install "breakingnews[gpu]" # + torch/transformers/peft, for inference
The base install is pure Python. Inference needs the [gpu] extra and a GPU with at least 24 GB; the adapter is fetched from the Hugging Face Hub on first use and cached, and the Llama-3.1-8B base is a further ~16 GB download. Built with Llama.
The workflow
# 1. Confirm the adapter is intact before trusting anything it produces.
breakingnews check-adapter sakshib3/Llama-3.1-breakingnews --revision v1
# 2. Transcripts -> boundaries.
breakingnews run sakshib3/Llama-3.1-breakingnews --revision v1 \
--input transcripts.jsonl --out predictions.jsonl
# 3. Boundaries -> one row per story.
breakingnews segments --transcripts transcripts.jsonl \
--predictions predictions.jsonl --out segments.jsonl --min-words 100
# 4. Rows -> whole documents again, byte for byte.
breakingnews merge --segments segments.jsonl --out rebuilt.jsonl
Two more commands help once you have output:
scorecompares a prediction file against your own gold annotations and reports precision, recall, F1 at three tolerances, plus Pk and WindowDiff. Every gold document counts, so a prediction file that is missing records is reported rather than quietly scoring higher.reconcilemaps segment ids from one run onto another by how much text they share, and tells you which ids carried over unchanged, which moved, and which were split or merged.
breakingnews score --predictions predictions.jsonl --gold annotations.jsonl
breakingnews reconcile --old run_a.jsonl --new run_b.jsonl
Everything except run and sweep works without a GPU.
In Python
from breakingnews import Segmenter, to_segments, merge_segments
seg = Segmenter.from_pretrained("sakshib3/Llama-3.1-breakingnews", revision="v1")
breaks = seg.segment(transcript) # [909, 1333, 2351, ...]
rows = to_segments(record_id, transcript, breaks, min_words=100)
merge_segments(rows) == (record_id, transcript) # byte for byte
Pin revision. Unpinned resolves to main, which moves.
What you get
One row per story. --minimal emits just the first four fields; offsets index transcript.split(), and schemas/ documents all four JSONL formats.
| field | |
|---|---|
record_id |
the parent broadcast, on every row |
segment_id |
{record_id}#{index:03d} |
text |
the story |
n_cuts |
boundaries found in the parent record; 0 means it was never cut |
word_start word_end char_start char_end n_words |
offsets |
Three properties the package holds to:
- Segmentation is a partition. Every character lands in exactly one story, so
mergereproduces the source byte for byte and refuses rather than guesses when it cannot. - Nothing is dropped silently.
--min-wordsflags short segments; a record that fails is named and the command exits non-zero. - Provenance survives.
record_idis a field on every row, never something you parse out of an id.
Re-running can shift a boundary by a few words, because batched bf16 generation is not bit-reproducible, so join two runs with reconcile, never on segment_id — that renumbers whenever a run finds a different number of stories.
Accuracy
τ (tau) is the decision threshold. For every window the model emits a probability that a story boundary is present in it; τ is the cut-off above which that window's boundaries are kept. τ = 0.010 here, selected on validation and applied unchanged to the held-out test set.
| split | docs | boundaries | tolerance | precision | recall | F1 |
|---|---|---|---|---|---|---|
| validation | 117 | 322 | ±25 w | 0.573 | 0.621 | 0.596 |
| validation | 117 | 322 | ±100 w | 0.728 | 0.789 | 0.757 |
| test | 20 | 64 | ±25 w | 0.593 | 0.797 | 0.680 |
| test | 20 | 64 | ±100 w | 0.663 | 0.891 | 0.760 |
Quote the validation row: it rests on 322 boundaries against the test split's 64, where the standard error on recall alone is ≈0.05. Precision is a lower bound, not an estimate — many scored false positives are real topic changes grouped into one thematic block, so do not compute a derived statistic that treats a false positive as clean error.
A prediction counts as correct if it lands within N words of a true boundary: ±25 asks "to within a sentence?", ±100 asks "did it find the seam at all?". Baselines on the same test set are 0.000 for predicting nothing and 0.062 for predicting N boundaries at uniform spacing.
τ is not a tuning knob. The confidences are saturated and bimodal — 49% of validation windows above 0.5, 38% below 0.001 — so any value in roughly [0.005, 0.5] gives the same answer, and it exists only to exclude τ = 0, where every window fires. This geometry has no high-precision regime: it cannot exceed precision 0.564 at any threshold, so if your use is sensitive to false boundaries the fix is a different geometry, not a different threshold.
Training
Trained on 998 annotated transcripts containing 2,829 boundaries, sampled from US TV news broadcasts 1992–2020 (CNN, FOX, MSNBC, ABC, CBS). The transcripts are licensed and are not distributed; the annotations are word offsets carrying no text, available for review on request. To train on your own corpus, see schemas/ for the input formats and scripts/model-training/ for the procedure.
License
The package is MIT. The model is a LoRA adapter on Llama-3.1-8B-Instruct and is governed by the Llama 3.1 Community License, whose terms pass through to anyone using the weights. Built with Llama.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file breakingnews-0.1.0.tar.gz.
File metadata
- Download URL: breakingnews-0.1.0.tar.gz
- Upload date:
- Size: 182.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dfff8e5b779054a5a875d3cc3652f3dd137f56de4a3fc29f3f2f1513df7cb3a1
|
|
| MD5 |
3f5dac0d44d197952fae9bcb76c738de
|
|
| BLAKE2b-256 |
93088d58b7fd022bd7a4c62bb41ec18202a45bb894dd0c0bb3a90be169a76627
|
Provenance
The following attestation bundles were made for breakingnews-0.1.0.tar.gz:
Publisher:
release.yml on sakshi-bhalla/breakingnews
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
breakingnews-0.1.0.tar.gz -
Subject digest:
dfff8e5b779054a5a875d3cc3652f3dd137f56de4a3fc29f3f2f1513df7cb3a1 - Sigstore transparency entry: 2363803540
- Sigstore integration time:
-
Permalink:
sakshi-bhalla/breakingnews@a23e7dedf5b420be9675da4dc80a160ba5bcdb0e -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/sakshi-bhalla
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@a23e7dedf5b420be9675da4dc80a160ba5bcdb0e -
Trigger Event:
push
-
Statement type:
File details
Details for the file breakingnews-0.1.0-py3-none-any.whl.
File metadata
- Download URL: breakingnews-0.1.0-py3-none-any.whl
- Upload date:
- Size: 50.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
876ded864b52c567bf4634590d330110f56c049fe77bab9a92486d601f925133
|
|
| MD5 |
fd59047cd68173f89f83629adcab15f8
|
|
| BLAKE2b-256 |
5956e9aec80cf5b73e56acb7d5a403972c8aeafea72d0e79b4f0ffa137fd9616
|
Provenance
The following attestation bundles were made for breakingnews-0.1.0-py3-none-any.whl:
Publisher:
release.yml on sakshi-bhalla/breakingnews
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
breakingnews-0.1.0-py3-none-any.whl -
Subject digest:
876ded864b52c567bf4634590d330110f56c049fe77bab9a92486d601f925133 - Sigstore transparency entry: 2363803653
- Sigstore integration time:
-
Permalink:
sakshi-bhalla/breakingnews@a23e7dedf5b420be9675da4dc80a160ba5bcdb0e -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/sakshi-bhalla
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@a23e7dedf5b420be9675da4dc80a160ba5bcdb0e -
Trigger Event:
push
-
Statement type: