Skip to main content

cjm-transcription-core

A frontend-agnostic core for the audio transcription workflow — composes isolated capability workers (audio conversion, VAD segmentation, batch transcription, persistence) into a headless pipeline, with a CLI as its first driver.

Modules

  • cjm_transcription_core.__init__
  • cjm_transcription_core.boundaries — Wall-clock-aware segment boundary computation: group VAD speech chunks into segments cut at silence-gap midpoints. Pure logic — no capability calls. Final home of the algorithm originally validated in cjm-transcription-audio-segment's AudioSegmentService.compute_segment_boundaries (that library is retired to cj-mills_deferred/).
  • cjm_transcription_core.candidates — Candidate (capability, MODEL)-instance enumeration for the comparison screen.
  • cjm_transcription_core.chunk — Chunk-grain re-runs, external landings and the runaway census (work item cf0b91d6; rulings 9ffce5f7 · 8a9b9639 · 910f3692): ONE landing, two producers — a local transcriber re-run of one AudioSegment's rendition, or an operator-pasted external transcript — lands a Transcript variant under the EXISTING AudioSegment with provenance, SUPERSEDES the prior variant for that (rendition, transcriber), journals the delta, and writes a DERIVED run manifest (parent_run_id + per-entry config_hash) the decomp consumes; the census names the chunks in need (degenerate-tail markers on new runs, oversized / implausible words-per-second text on old ones, extreme two-transcriber disagreement) and never counts a superseded variant as live.
  • cjm_transcription_core.cli — The CLI driver — the workflow core's first (and currently only) frontend.
  • cjm_transcription_core.curation — Collection curation vocabulary (hub v0, e5849229): the journaled update/delete
  • cjm_transcription_core.emission — Graph-root emission (CR-18 revolution 2): a completed source EMITS Source -> AudioSegment -> Transcript into the shared context graph — the graph BEGINS at transcription (where-graph-begins resolution: ingestion is the first EXTENDER that plants the root). Deterministic identity tuples make emission idempotent: re-runs (cache hits included) collide into verified no-ops instead of duplicating roots (the E13 hazard, relocated into graph creation and discharged).
  • cjm_transcription_core.launch — The shared launch surface every transcription shell drives through: the
  • cjm_transcription_core.models — Data shapes for the transcription pipeline: run configuration + the run-manifest result containers. The run manifest is the pipeline's durable output record: which sources were processed, how they were segmented, and where each segment's transcription landed (capability data DBs remain the authoritative text store; the manifest records the run's shape + provenance pointers). It is a deliberate proto-bundle — the CR-20 provenance-bundle infrastructure is expected to absorb/replace it.
  • cjm_transcription_core.pipeline — The headless transcription pipeline: VAD analysis -> boundary computation -> segment cutting -> per-segment model-input conversion -> transcription, composed over capability workers via the substrate's JobQueue. Between-stage outputs are threaded manually (run job -> read result -> submit next); the per-segment fan-out rides a CR-16 ports Composition with OutputRef bindings (this module was the real-world consumer of the original submit_sequence piping gap — pass-2 evidence in claude-docs/pass-2-evidence.md). HITL approval seams use the cheapest viable form (log + optional CLI prompt) per the cores-cluster guard-rails; each seam carries its 5-field HITL-assist annotation in its docstring.
  • cjm_transcription_core.probe — Per-segment comparison probe: transcribe ONE VAD-cut segment across every
  • cjm_transcription_core.results — Past-run results for the setup TUI: the core's own runs/*.json manifests read
  • cjm_transcription_core.sources — Source-selection state for the picker stage: a keyboard file browser plus the
  • cjm_transcription_core.state — Sidecar TUI state: last-used run settings persisted across sessions (the
  • claude-docs.mishomed_scan — 96edc646 population scan: flag FA words aligning outside every chunk of a spine.

API

cjm_transcription_core.boundaries

  • compute_segment_boundaries function — Group VAD chunks into segments cut at silence-gap midpoints.

cjm_transcription_core.candidates

  • candidate_directives function — Expand every installed transcription capability into its candidate space.
  • discover_capability function — Pick a DEFAULT capability for a role by surface match.
  • instance_id_for function — Derive an addressable instance id for a non-default (capability, MODEL) pick.
  • manifests_with_method function — Enumerate installed capabilities whose structural surface lists method.
  • model_axis function — Find a capability's MODEL config axis in its config_schema.
  • spec_string function — Render a load directive back to the core CLI's --transcriber grammar.
  • transcription_manifests function — Enumerate installed transcription capabilities from their manifest files.

cjm_transcription_core.chunk

  • apply_chunk_update function — Replace ONE chunk's entry for ONE transcriber in the derived manifest.
  • build_chunk_landing function — Build the landing payload for ONE chunk (pure; no capability calls).
  • census_rows function — The runaway census (pure): which LIVE chunk variants need a better transcription.
  • chunks_from_census function — Map census rows onto the manifest's chunks by (Source id, segment index) — the
  • derive_manifest function — Start a DERIVED manifest (ruling 910f3692 (1)): a copy of the parent under a NEW
  • fetch_transcript_rows function — Pull every Transcript's census inputs through the graph capability's marked
  • flagged_chunks function — The inspection lane's jump index: which chunks of a run carry a flagged variant,
  • is_external_transcriber function — Whether a transcriber name is an external landing's (the /manual marker).
  • land_chunk_transcript function — Land one chunk's variant through the task channel and journal the delta —
  • load_run_manifest function — Load a transcription-core run manifest (${WS}/ recorded paths resolve at load,
  • prior_config_hash function — The config hash of the variant this transcriber currently has on the chunk: a
  • prompt_hash_of function — Hash a prompt template — the prompt is DATA (f304d31d) and its hash rides the
  • render_escalation_prompt function — Render the escalation prompt WITH CONTEXT for one chunk (ruling 8a9b9639 (3):
  • rows_from_manifest function — Census inputs from a run manifest alone (no graph): the qt inspection lane's
  • save_manifest function — Write the derived manifest (the same recording contract as RunManifest.save).
  • select_chunks function — Resolve a chunk selection against the manifest (pure).
  • summarize_census function — Per-collection roll-up of the census (the closing evidence for 56a802b3 is a
  • text_shape function — Census the newline shape of a pasted external transcript (finding efe88f17;
  • wordwrap_warning function — The landing-time warning for a wordwrap-shaped paste (efe88f17 (2)): the

cjm_transcription_core.cli

  • add_transcript_command function — Execute add-transcript (cf0b91d6 part 2; ruling 9ffce5f7 (1)): land an
  • bind_source_dates_command function — Execute bind-source-dates (ruling de9c4cda (H7)). Two shapes: (1) --collection-id
  • bind_source_urls_command function — Execute bind-source-urls: join a collection's member Sources to a playlist
  • build_parser function — Build the CLI parser (subcommands: run).
  • declare_structure_command function — Execute declare-structure: read a structure-map document and land it
  • expand_sources function — Expand CLI source arguments into the ordered media-file list for a run.
  • expand_sources_with_collections function — Expand CLI sources AND keep the folder-source gesture as collection
  • load_capabilities function — Discover manifests + load each requested capability.
  • main function — CLI entry point (console script: cjm-transcription-core).
  • parse_config_overrides function — Parse repeatable KEY=VALUE config overrides (--transcriber-config).
  • parse_max_concurrent function — Parse repeatable --max-concurrent NAME=N values into a per-capability cap map.
  • parse_transcriber_spec function — Parse one --transcriber spec into a (capability, MODEL)-instance load directive.
  • probe_video_metadata_command function — Execute probe-video-metadata: ask yt-dlp for each playlist row's PER-VIDEO metadata
  • reference_command function — Execute add-reference / retract-reference: attach or retract a human-added
  • rerun_chunk_command function — Execute rerun-chunk (cf0b91d6 part 1; ruling 8a9b9639 chunk-targeted, never wholesale).
  • retire_collection_command function — Execute retire-collection (ruling a7617bd4, item eaefebd2): resolve the
  • run_command function — Execute the run subcommand: full pipeline over the given audio files.
  • runaway_census_command function — Execute runaway-census (cf0b91d6 part 3): the LIVE chunk variants in need of a

cjm_transcription_core.curation

  • add_reference function — Attach a HUMAN-ADDED RESOURCE LINK to a Source as a Reference NODE (ruling
  • apply_curation function — Replay one collection-curation op: deletes -> updates -> wires.
  • bind_source_dates function — Bind each Source's DATES — published_at (the public upload, an exact ISO day) and
  • bind_source_urls function — Bind each Source's PUBLIC URL — the time-addressable watch page a rendering links
  • collection_member_props function — A collection's member Sources WITH the named properties (PART_OF edges; unordered
  • collection_members function — A collection's member Sources (PART_OF edges; unordered by design —
  • collection_order function — Walk the materialized order, when one exists (typed EdgeQuery reads —
  • confirm_collection function — Discharge a proposed collection's flag (ae3464fc: the explicit human
  • curation_replay_handlers function — The curation verb's replay registration (unioned into
  • date_bindings_from_video_metadata function — Pure: join a collection's member Sources to per-video metadata rows BY VIDEO ID —
  • declare_structure function — Declare a SOURCE STRUCTURE MAP: the WORK's own part/chapter structure
  • file_sources function — File existing Sources into a collection (create-or-attach; the hub's
  • holding_collections function — The Collections holding a Source — the inverse of collection_members,
  • journal_curation function — Apply one curation act and journal it as a collection-curation op.
  • list_collections function — Enumerate the graph's Collection nodes (the hub's grouping corpus).
  • live_collections function — Filter retired collections out of a listing (pure; the pickers' default view).
  • refile_members function — Move members between collections (the Supernova carve-out: select
  • rename_collection function — Rename a collection — which IS merge when the new title already exists.
  • retire_collection function — Retire a Collection as a journaled FACT (ruling a7617bd4, item eaefebd2): status
  • retract_reference function — Retract a Reference node — the compensating act for add_reference (the node
  • set_collection_order function — Materialize (or repair) a collection's order — the curation op ae3464fc
  • sibling_sources function — A Source's collection siblings: the members of every LIVE collection
  • structure_entries_from_map function — Normalize a structure-map document into declare_structure entries.
  • url_bindings_from_playlist function — Pure: join a collection's member Sources to a playlist's rows BY TITLE and return

cjm_transcription_core.emission

  • build_collection_emission function — Build the Collection layer payload for one declaration (pure; no
  • build_source_emission function — Build the graph-root payload for one source (pure; no capability calls).
  • emit_collections_graph function — Idempotently emit the run's collection declarations (verb
  • emit_source_graph function — Idempotently emit one source's graph root through the task channel.
  • transcription_replay_handlers function — The transcription core's replay vocabulary (DEC 426658f1, replay stays DOMAIN-OWNED).

cjm_transcription_core.launch

  • build_parser function — The TUI driver's argument surface (setup options + core-run passthrough).
  • hand_off function — The shared driver tail: persist the confirmed choices, print the
  • plan_argv function — Render a confirmed plan as headless core-CLI argv.
  • resolve_settings function — Resolve the run-setup settings every shell shares (flags > persisted

cjm_transcription_core.models

  • CollectionDecl class — A collection declaration riding a run (ae3464fc: the folder-source
  • PipelineConfig class — Configuration for one transcription pipeline run.
  • RunManifest class — Durable record of one pipeline run (proto-bundle; see CR-20).
  • SegmentRecord class — One segment of a source audio file, with per-transcriber transcripts.
  • SourceResult class — Pipeline result for one source audio file.
  • new_run_id function — Generate a unique, sortable run id.

cjm_transcription_core.pipeline

  • acquire_speaker_turns function — Diarize the full source and persist the source-keyed turns artifact.
  • analyze_vad function — Run VAD analysis on one model-ready audio file (task channel: vad/detect_speech).
  • build_segment_composition function — Build the per-source fan-out composition: N independent [preprocess→]convert→(T× transcribe) pipes.
  • collect_capability_info function — Record capability identity + data-DB pointers for the run manifest (provenance).
  • confirm_seam function — HITL approval seam in its cheapest viable form (log + optional CLI prompt).
  • convert_for_vad function — Convert a source to MODEL-READY audio for VAD via the ffmpeg convert action.
  • cut_segments function — Cut the source audio at the computed boundaries via ffmpeg segment_audio.
  • normalize_vad_result function — Normalize a typed VAD result into sorted speech chunks + the reported duration.
  • probe_duration function — Probe a media file's duration via the ffmpeg capability's get_info action.
  • records_from_composition function — Fold a completed segment composition back into SegmentRecords.
  • run_pipeline function — Run the transcription pipeline over the given sources, in order.
  • run_source function — Run the full pipeline for one source: VAD → boundaries → cut → [preprocess →] convert → transcribe.
  • submit_and_wait function — Submit one capability job, wait for it, and return its result (raise on failure).
  • tier1_segment_checks function — Tier-1 deterministic pre-filters for the boundary-review seam (no AI).
  • tier1_transcript_checks function — Tier-1 deterministic pre-filters for the transcript-review seam (no AI).

cjm_transcription_core.probe

  • SegmentProbe class — One source's cut segments + cached per-segment comparison results.

cjm_transcription_core.results

  • RunIndex class — runs/*.json manifests loaded newest-first + the lookups the TUI paints from.

cjm_transcription_core.sources

  • CollectionField class — Pre-run collection state for the sources stage (ae3464fc: the actor
  • SourceBrowser class — Keyboard file-browser + ordered selection state for the sources stage.

cjm_transcription_core.state

  • load_state function — Read this project's persisted TUI state.
  • save_state function — Merge updates into the persisted state and write it back (best-effort:
  • state_path function — Where this project's TUI state lives.

claude-docs.mishomed_scan

  • scan function

Dependencies

Depends on: cjm-capability-primitives, cjm-context-graph-layer, cjm-context-graph-primitives, cjm-substrate, cjm-transcript-graph-schema, cjm-transcription-adapter-interface Used by: cjm-transcript-correction-qt, cjm-transcript-decomp-core, cjm-transcription-qt, cjm-workflow-hub-qt

Release files for cjm-transcription-core 0.0.22

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cjm-transcription-core 0.0.22
File Size Uploaded
cjm_transcription_core-0.0.22.tar.gz 118.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cjm-transcription-core 0.0.22
File Interpreter ABI Platform
cjm_transcription_core-0.0.22-py3-none-any.whl Python 3 none any Details

Total release size: 212.3 kB

Release files / cjm_transcription_core-0.0.22.tar.gz

Download URL cjm_transcription_core-0.0.22.tar.gz
Size 118.4 kB
Tags Source
SHA-256 checksum
How to use checksums
8b0a8806df98c587f788f7e789513cabccf1a1901ab5e720cb389cc1e3161841
BLAKE2b-256 checksum
How to use checksums
34ab3f7d99e44aa827f9ca34a3f40b3531480df34a86bc1e7dc6106ab449dc01
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release files / cjm_transcription_core-0.0.22-py3-none-any.whl

Download URL cjm_transcription_core-0.0.22-py3-none-any.whl
Size 93.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2512735f60754c61633a0e9d61460d85fd4112052cfc068bcd37b6450039cc5a
BLAKE2b-256 checksum
How to use checksums
11d7e0935459ccbb72153c17e8d1c903146941037e436bb8170a50eb3ba30c12
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release history Release notifications | RSS feed

This release

0.0.22 This release

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page