A frontend-agnostic core for the transcript decomposition workflow — composes isolated capability workers (forced alignment, VAD, graph storage) into a headless pipeline that decomposes transcription run manifests into a VAD-aligned context-graph spine with traceable provenance, and a CLI as its first driver.
Project description
cjm-transcript-decomp-core
A frontend-agnostic core for the transcript decomposition workflow — composes isolated capability workers (forced alignment, VAD, graph storage) into a headless pipeline that decomposes transcription run manifests into a VAD-aligned context-graph spine with traceable provenance, and a CLI as its first driver.
Modules
cjm_transcript_decomp_core.alignment— Pure forced-alignment logic (no capability calls): map FA words back to character spans in the original text, assign words to VAD chunks by timestamp, and build one text segment per VAD chunk. Extracted from the page-centric ForcedAlignmentService (Tier-1 logic).cjm_transcript_decomp_core.cli— The CLI driver — the decomposition core's first (and currently only) frontend.cjm_transcript_decomp_core.graph— Graph-spine EXTENSION + skeptical-lens verification (stage 5, CR-18 revolution 2). Decomp no longer creates a Document: it RECOMPUTES the transcription-emitted root's deterministic node ids from the consumed manifest (no search), verifies the root exists, and attaches the fine Segment spine under the existing AudioSegment nodes — PART_OF to the owning rendition, STARTS_WITH per rendition (the coarse-seam jump anchor), source-wide NEXT. Each Segment carries the audio TimeSlice ref plus per-transcriber CharSlice refs into the Transcript nodes (the D4/P10 framing, finally expressible). Commit goes through the layer's idempotent extend_graph.cjm_transcript_decomp_core.models— Lean data shapes for the transcript-decomposition pipeline: in-core mirrors of the forced-alignment / VAD / text DTOs (no FastHTML deps), run configuration, the committed graph-segment carrier, and the decomposition run manifest (proto-bundle).cjm_transcript_decomp_core.pipeline— The headless decomposition pipeline (stage 5: decomp is an EXTENDER). Load a transcription run manifest, verify the transcription-emitted graph root exists (the graph begins at transcription), then per source per pipeline-segment run VAD + per-transcriber forced alignment, build one aligned segment per VAD chunk with per-transcriber text variants, and attach the fine spine under the existing AudioSegment nodes via the layer's idempotent extend_graph — with HITL approval seams between alignment, commit, and the next source.
API
cjm_transcript_decomp_core.alignment
assign_words_to_chunksfunction — Assign each FA word to a VAD chunk by timestamp overlap.build_segments_from_alignmentfunction — Build a TextSegment per VAD chunk by grouping words by chunk assignment.map_fa_words_to_textfunction — Map forced-alignment words back to character spans in the original text.tier1_alignment_checksfunction — Tier-1 deterministic pre-filters for the alignment-review seam (no AI).
cjm_transcript_decomp_core.cli
build_parserfunction — Build the CLI parser (subcommands: run).load_capabilitiesfunction — Discover manifests + load each requested capability (default instance).mainfunction — CLI entry point (console script:cjm-transcript-decomp-core).run_commandfunction — Execute therunsubcommand: extend transcription-run manifest(s) with the fine spine.
cjm_transcript_decomp_core.graph
SourceVerificationclass — Skeptical-lens verification of one Source's fine-spine extension under abuild_extension_payloadfunction — Build the fine-spine EXTENSION payload (pure; no capability calls).resolve_root_idsfunction — Recompute the transcription-emitted root node ids from manifest data.verify_sourcefunction — Verify a Source's committed extension via server-side AGGREGATES (D13/D19).
cjm_transcript_decomp_core.models
DecompConfigclass — Configuration for one transcript-decomposition run.DecompManifestclass — Durable record of one decomposition run (proto-bundle; see CR-20).DecompSegmentclass — One fine spine segment (stage 5: shared audio-side skeleton + per-transcriber variants).DecompSourceRecordclass — Record of one Source whose fine spine this run committed (stage 5:FAWordclass — One word-level forced-alignment result (segment-local times).SegmentVariantclass — One transcriber's text + char range for one fine segment (stage 5).TextSegmentclass — A text segment produced by alignment, before graph commit.VADChunkclass — A voice-activity time range within one pipeline segment (segment-local).new_run_idfunction — Generate a unique, sortable decomposition run id.
cjm_transcript_decomp_core.pipeline
build_alignment_compositionfunction — Build the whole-source M×(VAD ∥ T×FA) composition (D8 fan-in, stage-5 variants).collect_capability_infofunction — Record capability identity + data-DB pointers for the run manifest (provenance).confirm_seamfunction — HITL approval seam in its cheapest viable form (log + optional CLI prompt).decomp_replay_handlersfunction — The decomp core's replay vocabulary (DEC 426658f1, replay stays DOMAIN-OWNED).decompose_sourcefunction — Decompose one source into aligned fine segments with per-transcriber variants.fa_words_from_resultfunction — Normalize a typed forced-alignment result into FA words (pure; stage 3).load_source_manifestfunction — Load + lightly validate a transcription-core run manifest.run_decompfunction — Extend every source in a transcription run manifest with its fine spine.submit_and_waitfunction — Submit one capability job, wait for it, and return its result (raise on failure).vad_chunks_from_resultfunction — Normalize a typed VAD result into segment-local VAD chunks.
Dependencies
Depends on: cjm-capability-primitives, cjm-context-graph-layer, cjm-context-graph-primitives, cjm-substrate, cjm-transcript-graph-schema
Used by: cjm-transcript-decomp-tui
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cjm_transcript_decomp_core-0.0.4.tar.gz.
File metadata
- Download URL: cjm_transcript_decomp_core-0.0.4.tar.gz
- Upload date:
- Size: 34.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
13c04b462244dd7d35dc114e14315e61e457d5da23ccca3b599d5c060e70ce35
|
|
| MD5 |
a0500920c693e82392e3b936aac089eb
|
|
| BLAKE2b-256 |
14b71f694f4f54575165c142f3e003b6f60db137fb9432403bf2cc5318457995
|
File details
Details for the file cjm_transcript_decomp_core-0.0.4-py3-none-any.whl.
File metadata
- Download URL: cjm_transcript_decomp_core-0.0.4-py3-none-any.whl
- Upload date:
- Size: 32.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
07ebea4feea1ae07719285471405267bdac4aa3c371852f06c81d060b6cab336
|
|
| MD5 |
ce1765f1dee7a960acd66dd60e151430
|
|
| BLAKE2b-256 |
a2b16f8f0ec48594a0987588b4ecfc73c6aaba963be7d278adca6a02615a6ec8
|