Skip to main content

cjm-transcript-decomp-core

A frontend-agnostic core for the transcript decomposition workflow — composes isolated capability workers (forced alignment, VAD, graph storage) into a headless pipeline that decomposes transcription run manifests into a VAD-aligned context-graph spine with traceable provenance, and a CLI as its first driver.

Modules

  • cjm_transcript_decomp_core.__init__
  • cjm_transcript_decomp_core.alignment — Pure forced-alignment logic (no capability calls): map FA words back to character spans in the original text, assign words to VAD chunks by timestamp, and build one text segment per VAD chunk. Extracted from the page-centric ForcedAlignmentService (Tier-1 logic).
  • cjm_transcript_decomp_core.cli — The CLI driver — the decomposition core's first (and currently only) frontend.
  • cjm_transcript_decomp_core.discovery — Capability-role discovery by manifest surface match — the journaling-by-
  • cjm_transcript_decomp_core.graph — Graph-spine EXTENSION + skeptical-lens verification (stage 5, CR-18 revolution 2). Decomp no longer creates a Document: it RECOMPUTES the transcription-emitted root's deterministic node ids from the consumed manifest (no search), verifies the root exists, and attaches the fine Segment spine under the existing AudioSegment nodes — PART_OF to the owning rendition, STARTS_WITH per rendition (the coarse-seam jump anchor), source-wide NEXT. Each Segment carries the audio TimeSlice ref plus per-transcriber CharSlice refs into the Transcript nodes (the D4/P10 framing, finally expressible). Commit goes through the layer's idempotent extend_graph.
  • cjm_transcript_decomp_core.launch — The shared launch surface every decomp shell drives through: the argument
  • cjm_transcript_decomp_core.models — Lean data shapes for the transcript-decomposition pipeline: in-core mirrors of the forced-alignment / VAD / text DTOs (no FastHTML deps), run configuration, the committed graph-segment carrier, and the decomposition run manifest (proto-bundle).
  • cjm_transcript_decomp_core.pipeline — The headless decomposition pipeline (stage 5: decomp is an EXTENDER). Load a transcription run manifest, verify the transcription-emitted graph root exists (the graph begins at transcription), then per source per pipeline-segment run VAD + per-transcriber forced alignment, build one aligned segment per VAD chunk with per-transcriber text variants, and attach the fine spine under the existing AudioSegment nodes via the layer's idempotent extend_graph — with HITL approval seams between alignment, commit, and the next source.
  • cjm_transcript_decomp_core.respine — Chunk RESPINE — re-derive ONE coarse chunk of a LIVE spine from a landed external transcript (work item 7a5e9c84; design ruling 0b4d5cfa, amendment 4a7ec4f8).
  • cjm_transcript_decomp_core.retire — Spine RETIREMENT + COMPACTION — safe removal of superseded decomposition spines (ruling a7617bd4, item eaefebd2).
  • cjm_transcript_decomp_core.runs — Run-manifest indexes for the decomp-batch TUI (work item 0ff6bf0f): the
  • cjm_transcript_decomp_core.segments — Fine-segment inspection for the decomp TUI (work item 166dd2b8, half a):
  • cjm_transcript_decomp_core.state — Sidecar TUI state: last-used batch settings persisted across sessions (the

API

cjm_transcript_decomp_core.alignment

  • assign_words_to_chunks function — Assign each FA word to a VAD chunk by timestamp overlap.
  • build_segments_from_alignment function — Build a TextSegment per VAD chunk by grouping words by chunk assignment.
  • carve_chunks_at_event_spans function — The event-carve stage (EVENT_SPLIT_POLICY, respine trial DEC 6cc10fb7):
  • collapse_newlines function — The stored-text form of a slice that still holds a newline: a retained
  • map_fa_words_to_text function — Map forced-alignment words back to character spans in the original text.
  • normalize_external_text function — Offset-preserving fold-input normalisation for EXTERNAL variants (finding
  • rescue_gap_words function — The word-rescue stage (WORD_RESCUE_POLICY, 96edc646 verdict bc7ece7b):
  • sentence_end_word_indices function — Map capability-delivered sentence boundaries onto FA words (B.5: the
  • split_chunks_at_sentence_gaps function — The sentence-split stage (SENTENCE_SPLIT_POLICY, DEC f1024568): refine the
  • tier1_alignment_checks function — Tier-1 deterministic pre-filters for the alignment-review seam (no AI).

cjm_transcript_decomp_core.cli

  • backfill_provenance_command function — Execute backfill-provenance: mint the Segment -> Transcript DERIVED_FROM
  • build_parser function — Build the CLI parser (subcommands: run).
  • load_capabilities function — Discover manifests + load each requested capability (default instance).
  • main function — CLI entry point (console script: cjm-transcript-decomp-core).
  • respine_chunk_command function — Execute respine-chunk (work item 7a5e9c84, ruling 0b4d5cfa; the machinery lives
  • run_command function — Execute the run subcommand: extend transcription-run manifest(s) with the fine spine.
  • spine_command function — The spine retirement + compaction verbs (ruling a7617bd4; the machinery lives in

cjm_transcript_decomp_core.discovery

  • discover_capability function — Pick a DEFAULT capability for a role by surface match.
  • manifests_with_method function — Enumerate installed capabilities whose structural surface lists method.

cjm_transcript_decomp_core.graph

  • SourceVerification class — Skeptical-lens verification of one Source's fine-spine extension under a
  • build_extension_payload function — Build the fine-spine EXTENSION payload (pure; no capability calls).
  • provenance_edges_from_segment_wire function — Derive a Segment's Transcript provenance edges from the node itself:
  • resolve_root_ids function — Recompute the transcription-emitted root node ids from manifest data.
  • segment_provenance_edges function — Text provenance as EDGES (finding 89b16be6, the references-must-be-edges
  • verify_source function — Verify a Source's committed extension via server-side AGGREGATES (D13/D19).

cjm_transcript_decomp_core.launch

  • batch_argv function — Render one hand-off group as headless decomp-core argv.
  • build_parser function — The TUI driver's argument surface (batch-setup options + core passthrough).
  • event_split_batch_error function — One propset carves ONE source: the core applies --event-propset to every
  • hand_off function — The shared driver tail: persist the confirmed choices, adopt the in-app
  • resolve_settings function — Resolve the batch-setup settings every shell shares (flags > persisted
  • resolve_split_flags function — Resolve the post-parse split-flag contract (pure; main() calls it

cjm_transcript_decomp_core.models

  • DecompConfig class — Configuration for one transcript-decomposition run.
  • DecompManifest class — Durable record of one decomposition run (proto-bundle; see CR-20).
  • DecompSegment class — One fine spine segment (stage 5: shared audio-side skeleton + per-transcriber variants).
  • DecompSourceRecord class — Record of one Source whose fine spine this run committed (stage 5:
  • FAWord class — One word-level forced-alignment result (segment-local times).
  • SegmentVariant class — One transcriber's text + char range for one fine segment (stage 5).
  • TextSegment class — A text segment produced by alignment, before graph commit.
  • VADChunk class — A voice-activity time range within one pipeline segment (segment-local).
  • new_run_id function — Generate a unique, sortable decomposition run id.

cjm_transcript_decomp_core.pipeline

  • build_alignment_composition function — Build the whole-source M×(VAD ∥ T×FA ∥ SEG) composition (D8 fan-in, stage-5 variants).
  • collect_capability_info function — Record capability identity + data-DB pointers for the run manifest (provenance).
  • compute_skeleton_hash function — Skeleton identity, pure (DEC f1024568 + 9241564f + 6cc10fb7).
  • confirm_seam function — HITL approval seam in its cheapest viable form (log + optional CLI prompt).
  • decomp_replay_handlers function — The decomp core's replay vocabulary (DEC 426658f1, replay stays DOMAIN-OWNED).
  • decompose_source function — Decompose one source into aligned fine segments with per-transcriber variants.
  • event_spans_from_propset function — Load a proposal set BY POINTER and select its carve spans (respine trial
  • fa_words_from_result function — Normalize a typed forced-alignment result into FA words (pure; stage 3).
  • load_source_manifest function — Load + lightly validate a transcription-core run manifest.
  • resolve_event_propsets function — Join each proposal set to ITS source by the set manifest's own source
  • run_decomp function — Extend every source in a transcription run manifest with its fine spine.
  • sentence_spans_from_result function — Normalize a typed sentence-segmentation result into char-span tuples (pure; B.5).
  • submit_and_wait function — Submit one capability job, wait for it, and return its result (raise on failure).
  • vad_chunks_from_result function — Normalize a typed VAD result into segment-local VAD chunks.

cjm_transcript_decomp_core.respine

  • bridge_edges function — Bridge NEXT from the previous live segment to the new first and from the new
  • build_respine_op function — ONE replayed property-update op (pure): superseded_by on the old segments, the
  • chunk_entry_for function — Resolve ONE coarse chunk of a source entry (pure): by manifest index, or by
  • classify_dependents function — Sort the chunk's dependents into the two transferable classes and the rest
  • decomp_config_from function — Rebuild the live spine's DecompConfig from its manifest snapshot (pure): the
  • describe_dependents function — The refusal / readout listing (pure).
  • find_decomp_manifest function — The decomp run that minted the live spine (pure): the newest manifest whose
  • live_chunk_segments function — The chunk's LIVE segments: PART_OF its rendition, on the spine, not superseded.
  • live_spine_index function — The whole live spine's (id, index) — the renumber plan's input.
  • plan_renumber function — The renumber plan (pure; 0b4d5cfa (4), the 'do it properly' ruling): the new
  • read_dependents function — What points at the old segments: every Correction with a CORRECTS edge into
  • render_chunk_prompt function — Mode one (--prompt): the chunk's escalation prompt WITH CONTEXT — no writes.
  • resolve_chunk_context function — Resolve the live spine -> its decomp manifest -> the transcription manifest it
  • respine_chunk function — Mode two (--text-file): land -> re-derive the chunk under the live policy ->
  • respine_op_id function — The op id superseded_by names: a function of (source, landed transcript).
  • source_entry_for function — The transcription manifest's entry for a Source, by recomputed identity (pure).
  • transfer_handler function — Discover the correction core's chunk-scoped transfer through the entry-point

cjm_transcript_decomp_core.retire

  • annotate_spines function — Mark each spine row with its retirement state (pure; rows copied).
  • apply_spine_fact function — Replay handler for spine-retire / spine-compaction / chunk-respine: property merges
  • compact_retired function — The compact act over every retired, not-yet-compacted spine of the given sources:
  • default_live_spine function — Which spine opens by default (pure). Preference is a DECLARED fact, never creation order:
  • dependents_free function
  • dependents_map function — Every spine's dependents in one raw read (see DEPENDENTS_SQL); a spine absent from
  • get_source function — The Source node as a dict (None when absent).
  • journal_spine_retire function — Apply + journal one retirement fact (write-side dual of apply_spine_fact).
  • list_sources function — Every Source (optionally one collection's members).
  • list_spines function — The source's coexisting spines, grouped by skeleton hash and annotated with the retirement map.
  • live_spines function
  • plan_retire function — Rules (a) + successor validation as a pure plan; the graph write is journal_spine_retire.
  • plan_superseded function — The batch candidate list: every LIVE spine that is not the source's default (rule (c)) —
  • resolve_source_id function — Resolve a human selector to ONE Source (refuses with the candidates).
  • resolve_spine function — Resolve a picker-style selector to exactly one spine row (pure); refuses with the roster.
  • retire_spine function — The whole retire act: list -> plan (rules a) -> dependents gate (rule b) -> journaled fact.
  • retired_spines function — The Source's retirement map (pure): {spine key: {reason, successor, actor, ts[, compacted]}}.
  • source_rendition_ids function — Every AudioRendition under a Source (all chains — retirement is per skeleton, not per chain).
  • spine_dependents function — Rule (b)'s evidence: what on the graph points at the spine's segments.
  • spine_fact_handlers function
  • spine_key function
  • spine_label function
  • spine_segment_ids function

cjm_transcript_decomp_core.runs

  • DecompIndex class — Decomp-core run manifests read back: coverage chips for the batch stage
  • PropsetIndex class — Proposal-set manifests under the workspace proposals/ dir — the model's
  • SourceRunIndex class — Transcription-core run manifests — the decomp workflow's SOURCES — plus
  • TrainingRunIndex class — Training-run manifests under the workspace training-runs/ dir — the
  • group_batches function — Fold an ordered batch selection into headless hand-off groups.
  • is_external_transcriber function — Whether a transcriber name is an external landing's (the /manual marker).

cjm_transcript_decomp_core.segments

  • SegmentStack class — One lazily-opened READ-ONLY graph seat, keyed by db path.
  • aseg_index_for function — Which coarse AudioSegment a source-coordinate time falls in (pure;
  • build_display function — Interleave gap markers into the paintable entry list (pure).
  • capability_config_schema function — A capability's config_schema off its installed manifest (json read, pure).
  • find_gaps function — Uncovered timestamp spans between committed segments (pure).
  • fmt_ts function — Source-coordinate timestamp for listing rows (pure).
  • locate_span function — Resolve a source-coordinate span onto its owning coarse WAV (pure).
  • ordered_rows function — Re-impose the manifest's spine order on fetched rows (pure).
  • predicted_rows function — Synthesize segment-shaped rows for a probe's predicted skeleton (pure).
  • probe_compare function — Compare a VAD probe against the committed skeleton (pure).
  • realign_rows function — Re-run the decomp pipeline's own text fold over a PROBE skeleton (pure).
  • split_predicted function — Run the decomp pipeline's own SENTENCE-SPLIT stage over a probe skeleton
  • vad_summary function — The decomp run's recorded VAD config, one status line (pure).

cjm_transcript_decomp_core.state

  • load_state function — Read this project's persisted TUI state.
  • save_state function — Merge updates into the persisted state and write it back (best-effort:
  • state_path function — Where this project's TUI state lives.

Dependencies

Depends on: cjm-capability-primitives, cjm-context-graph-layer, cjm-context-graph-primitives, cjm-substrate, cjm-transcript-graph-schema, cjm-transcription-core, pyyaml Used by: cjm-transcript-correction-qt, cjm-transcript-decomp-qt

Release files for cjm-transcript-decomp-core 0.0.20

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cjm-transcript-decomp-core 0.0.20
File Size Uploaded
cjm_transcript_decomp_core-0.0.20.tar.gz 133.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cjm-transcript-decomp-core 0.0.20
File Interpreter ABI Platform
cjm_transcript_decomp_core-0.0.20-py3-none-any.whl Python 3 none any Details

Total release size: 238.9 kB

Release files / cjm_transcript_decomp_core-0.0.20.tar.gz

Download URL cjm_transcript_decomp_core-0.0.20.tar.gz
Size 133.4 kB
Tags Source
SHA-256 checksum
How to use checksums
c3518801ac48c141e9523381eda3573eedbbc35ed646ea31ab8e9443c2a762dc
BLAKE2b-256 checksum
How to use checksums
e4b859c3ec0daa60b5e57d361f741748b342ca1bbcfd2ba4d437a65074138a94
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release files / cjm_transcript_decomp_core-0.0.20-py3-none-any.whl

Download URL cjm_transcript_decomp_core-0.0.20-py3-none-any.whl
Size 105.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c29bd2c3e67411a75156e001741685687ccdd631b54328f33244b53c815fd722
BLAKE2b-256 checksum
How to use checksums
26b0cc194d55db9a22d23a3591e0e09d1cf7f2006b4373fa2ea5814c91394bad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release history Release notifications | RSS feed

This release

0.0.20 This release

2 release files

0.0.9

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page