Skip to main content

A frontend-agnostic core for the transcript decomposition workflow — composes isolated capability workers (forced alignment, VAD, graph storage) into a headless pipeline that decomposes transcription run manifests into a VAD-aligned context-graph spine with traceable provenance, and a CLI as its first driver.

Project description

cjm-transcript-decomp-core

A frontend-agnostic core for the transcript decomposition workflow — composes isolated capability workers (forced alignment, VAD, graph storage) into a headless pipeline that decomposes transcription run manifests into a VAD-aligned context-graph spine with traceable provenance, and a CLI as its first driver.

Modules

  • cjm_transcript_decomp_core.alignment — Pure forced-alignment logic (no capability calls): map FA words back to character spans in the original text, assign words to VAD chunks by timestamp, and build one text segment per VAD chunk. Extracted from the page-centric ForcedAlignmentService (Tier-1 logic).
  • cjm_transcript_decomp_core.cli — The CLI driver — the decomposition core's first (and currently only) frontend.
  • cjm_transcript_decomp_core.graph — Graph-spine EXTENSION + skeptical-lens verification (stage 5, CR-18 revolution 2). Decomp no longer creates a Document: it RECOMPUTES the transcription-emitted root's deterministic node ids from the consumed manifest (no search), verifies the root exists, and attaches the fine Segment spine under the existing AudioSegment nodes — PART_OF to the owning rendition, STARTS_WITH per rendition (the coarse-seam jump anchor), source-wide NEXT. Each Segment carries the audio TimeSlice ref plus per-transcriber CharSlice refs into the Transcript nodes (the D4/P10 framing, finally expressible). Commit goes through the layer's idempotent extend_graph.
  • cjm_transcript_decomp_core.models — Lean data shapes for the transcript-decomposition pipeline: in-core mirrors of the forced-alignment / VAD / text DTOs (no FastHTML deps), run configuration, the committed graph-segment carrier, and the decomposition run manifest (proto-bundle).
  • cjm_transcript_decomp_core.pipeline — The headless decomposition pipeline (stage 5: decomp is an EXTENDER). Load a transcription run manifest, verify the transcription-emitted graph root exists (the graph begins at transcription), then per source per pipeline-segment run VAD + per-transcriber forced alignment, build one aligned segment per VAD chunk with per-transcriber text variants, and attach the fine spine under the existing AudioSegment nodes via the layer's idempotent extend_graph — with HITL approval seams between alignment, commit, and the next source.

API

cjm_transcript_decomp_core.alignment

  • assign_words_to_chunks function — Assign each FA word to a VAD chunk by timestamp overlap.
  • build_segments_from_alignment function — Build a TextSegment per VAD chunk by grouping words by chunk assignment.
  • map_fa_words_to_text function — Map forced-alignment words back to character spans in the original text.
  • tier1_alignment_checks function — Tier-1 deterministic pre-filters for the alignment-review seam (no AI).

cjm_transcript_decomp_core.cli

  • build_parser function — Build the CLI parser (subcommands: run).
  • load_capabilities function — Discover manifests + load each requested capability (default instance).
  • main function — CLI entry point (console script: cjm-transcript-decomp-core).
  • run_command function — Execute the run subcommand: extend transcription-run manifest(s) with the fine spine.

cjm_transcript_decomp_core.graph

  • SourceVerification class — Skeptical-lens verification of one Source's fine-spine extension under a
  • build_extension_payload function — Build the fine-spine EXTENSION payload (pure; no capability calls).
  • resolve_root_ids function — Recompute the transcription-emitted root node ids from manifest data.
  • verify_source function — Verify a Source's committed extension via server-side AGGREGATES (D13/D19).

cjm_transcript_decomp_core.models

  • DecompConfig class — Configuration for one transcript-decomposition run.
  • DecompManifest class — Durable record of one decomposition run (proto-bundle; see CR-20).
  • DecompSegment class — One fine spine segment (stage 5: shared audio-side skeleton + per-transcriber variants).
  • DecompSourceRecord class — Record of one Source whose fine spine this run committed (stage 5:
  • FAWord class — One word-level forced-alignment result (segment-local times).
  • SegmentVariant class — One transcriber's text + char range for one fine segment (stage 5).
  • TextSegment class — A text segment produced by alignment, before graph commit.
  • VADChunk class — A voice-activity time range within one pipeline segment (segment-local).
  • new_run_id function — Generate a unique, sortable decomposition run id.

cjm_transcript_decomp_core.pipeline

  • build_alignment_composition function — Build the whole-source M×(VAD ∥ T×FA) composition (D8 fan-in, stage-5 variants).
  • collect_capability_info function — Record capability identity + data-DB pointers for the run manifest (provenance).
  • confirm_seam function — HITL approval seam in its cheapest viable form (log + optional CLI prompt).
  • decomp_replay_handlers function — The decomp core's replay vocabulary (DEC 426658f1, replay stays DOMAIN-OWNED).
  • decompose_source function — Decompose one source into aligned fine segments with per-transcriber variants.
  • fa_words_from_result function — Normalize a typed forced-alignment result into FA words (pure; stage 3).
  • load_source_manifest function — Load + lightly validate a transcription-core run manifest.
  • run_decomp function — Extend every source in a transcription run manifest with its fine spine.
  • submit_and_wait function — Submit one capability job, wait for it, and return its result (raise on failure).
  • vad_chunks_from_result function — Normalize a typed VAD result into segment-local VAD chunks.

Dependencies

Depends on: cjm-capability-primitives, cjm-context-graph-layer, cjm-context-graph-primitives, cjm-substrate, cjm-transcript-graph-schema Used by: cjm-transcript-decomp-tui

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cjm_transcript_decomp_core-0.0.4.tar.gz (34.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cjm_transcript_decomp_core-0.0.4-py3-none-any.whl (32.2 kB view details)

Uploaded Python 3

File details

Details for the file cjm_transcript_decomp_core-0.0.4.tar.gz.

File metadata

File hashes

Hashes for cjm_transcript_decomp_core-0.0.4.tar.gz
Algorithm Hash digest
SHA256 13c04b462244dd7d35dc114e14315e61e457d5da23ccca3b599d5c060e70ce35
MD5 a0500920c693e82392e3b936aac089eb
BLAKE2b-256 14b71f694f4f54575165c142f3e003b6f60db137fb9432403bf2cc5318457995

See more details on using hashes here.

File details

Details for the file cjm_transcript_decomp_core-0.0.4-py3-none-any.whl.

File metadata

File hashes

Hashes for cjm_transcript_decomp_core-0.0.4-py3-none-any.whl
Algorithm Hash digest
SHA256 07ebea4feea1ae07719285471405267bdac4aa3c371852f06c81d060b6cab336
MD5 ce1765f1dee7a960acd66dd60e151430
BLAKE2b-256 a2b16f8f0ec48594a0987588b4ecfc73c6aaba963be7d278adca6a02615a6ec8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page