Skip to main content

cjm-transcription-core

A frontend-agnostic core for the audio transcription workflow — composes isolated capability workers (audio conversion, VAD segmentation, batch transcription, persistence) into a headless pipeline, with a CLI as its first driver.

Modules

  • cjm_transcription_core.boundaries — Wall-clock-aware segment boundary computation: group VAD speech chunks into segments cut at silence-gap midpoints. Pure logic — no capability calls. Final home of the algorithm originally validated in cjm-transcription-audio-segment's AudioSegmentService.compute_segment_boundaries (that library is retired to cj-mills_deferred/).
  • cjm_transcription_core.cli — The CLI driver — the workflow core's first (and currently only) frontend.
  • cjm_transcription_core.curation — Collection curation vocabulary (hub v0, e5849229): the journaled update/delete
  • cjm_transcription_core.emission — Graph-root emission (CR-18 revolution 2): a completed source EMITS Source -> AudioSegment -> Transcript into the shared context graph — the graph BEGINS at transcription (where-graph-begins resolution: ingestion is the first EXTENDER that plants the root). Deterministic identity tuples make emission idempotent: re-runs (cache hits included) collide into verified no-ops instead of duplicating roots (the E13 hazard, relocated into graph creation and discharged).
  • cjm_transcription_core.models — Data shapes for the transcription pipeline: run configuration + the run-manifest result containers. The run manifest is the pipeline's durable output record: which sources were processed, how they were segmented, and where each segment's transcription landed (capability data DBs remain the authoritative text store; the manifest records the run's shape + provenance pointers). It is a deliberate proto-bundle — the CR-20 provenance-bundle infrastructure is expected to absorb/replace it.
  • cjm_transcription_core.pipeline — The headless transcription pipeline: VAD analysis -> boundary computation -> segment cutting -> per-segment model-input conversion -> transcription, composed over capability workers via the substrate's JobQueue. Between-stage outputs are threaded manually (run job -> read result -> submit next); the per-segment fan-out rides a CR-16 ports Composition with OutputRef bindings (this module was the real-world consumer of the original submit_sequence piping gap — pass-2 evidence in claude-docs/pass-2-evidence.md). HITL approval seams use the cheapest viable form (log + optional CLI prompt) per the cores-cluster guard-rails; each seam carries its 5-field HITL-assist annotation in its docstring.

API

cjm_transcription_core.boundaries

  • compute_segment_boundaries function — Group VAD chunks into segments cut at silence-gap midpoints.

cjm_transcription_core.cli

  • build_parser function — Build the CLI parser (subcommands: run).
  • expand_sources function — Expand CLI source arguments into the ordered media-file list for a run.
  • expand_sources_with_collections function — Expand CLI sources AND keep the folder-source gesture as collection
  • load_capabilities function — Discover manifests + load each requested capability.
  • main function — CLI entry point (console script: cjm-transcription-core).
  • parse_max_concurrent function — Parse repeatable --max-concurrent NAME=N values into a per-capability cap map.
  • parse_transcriber_spec function — Parse one --transcriber spec into a (capability, MODEL)-instance load directive.
  • run_command function — Execute the run subcommand: full pipeline over the given audio files.

cjm_transcription_core.curation

  • apply_curation function — Replay one collection-curation op: deletes -> updates -> wires.
  • collection_members function — A collection's member Sources (PART_OF edges; unordered by design —
  • collection_order function — Walk the materialized order, when one exists (typed EdgeQuery reads —
  • confirm_collection function — Discharge a proposed collection's flag (ae3464fc: the explicit human
  • curation_replay_handlers function — The curation verb's replay registration (unioned into
  • file_sources function — File existing Sources into a collection (create-or-attach; the hub's
  • journal_curation function — Apply one curation act and journal it as a collection-curation op.
  • list_collections function — Enumerate the graph's Collection nodes (the hub's grouping corpus).
  • refile_members function — Move members between collections (the Supernova carve-out: select
  • rename_collection function — Rename a collection — which IS merge when the new title already exists.
  • set_collection_order function — Materialize (or repair) a collection's order — the curation op ae3464fc

cjm_transcription_core.emission

  • build_collection_emission function — Build the Collection layer payload for one declaration (pure; no
  • build_source_emission function — Build the graph-root payload for one source (pure; no capability calls).
  • emit_collections_graph function — Idempotently emit the run's collection declarations (verb
  • emit_source_graph function — Idempotently emit one source's graph root through the task channel.
  • transcription_replay_handlers function — The transcription core's replay vocabulary (DEC 426658f1, replay stays DOMAIN-OWNED).

cjm_transcription_core.models

  • CollectionDecl class — A collection declaration riding a run (ae3464fc: the folder-source
  • PipelineConfig class — Configuration for one transcription pipeline run.
  • RunManifest class — Durable record of one pipeline run (proto-bundle; see CR-20).
  • SegmentRecord class — One segment of a source audio file, with per-transcriber transcripts.
  • SourceResult class — Pipeline result for one source audio file.
  • new_run_id function — Generate a unique, sortable run id.

cjm_transcription_core.pipeline

  • acquire_speaker_turns function — Diarize the full source and persist the source-keyed turns artifact.
  • analyze_vad function — Run VAD analysis on one model-ready audio file (task channel: vad/detect_speech).
  • build_segment_composition function — Build the per-source fan-out composition: N independent [preprocess→]convert→(T× transcribe) pipes.
  • collect_capability_info function — Record capability identity + data-DB pointers for the run manifest (provenance).
  • confirm_seam function — HITL approval seam in its cheapest viable form (log + optional CLI prompt).
  • convert_for_vad function — Convert a source to MODEL-READY audio for VAD via the ffmpeg convert action.
  • cut_segments function — Cut the source audio at the computed boundaries via ffmpeg segment_audio.
  • normalize_vad_result function — Normalize a typed VAD result into sorted speech chunks + the reported duration.
  • probe_duration function — Probe a media file's duration via the ffmpeg capability's get_info action.
  • records_from_composition function — Fold a completed segment composition back into SegmentRecords.
  • run_pipeline function — Run the transcription pipeline over the given sources, in order.
  • run_source function — Run the full pipeline for one source: VAD → boundaries → cut → [preprocess →] convert → transcribe.
  • submit_and_wait function — Submit one capability job, wait for it, and return its result (raise on failure).
  • tier1_segment_checks function — Tier-1 deterministic pre-filters for the boundary-review seam (no AI).
  • tier1_transcript_checks function — Tier-1 deterministic pre-filters for the transcript-review seam (no AI).

Dependencies

Depends on: cjm-capability-primitives, cjm-context-graph-layer, cjm-context-graph-primitives, cjm-substrate, cjm-transcript-graph-schema, cjm-transcription-adapter-interface Used by: cjm-transcription-tui, cjm-workflow-hub-tui

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cjm_transcription_core-0.0.7.tar.gz (50.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cjm_transcription_core-0.0.7-py3-none-any.whl (43.4 kB view details)

Uploaded Python 3

File details

Details for the file cjm_transcription_core-0.0.7.tar.gz.

File metadata

  • Download URL: cjm_transcription_core-0.0.7.tar.gz
  • Upload date:
  • Size: 50.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.12

File hashes

Hashes for cjm_transcription_core-0.0.7.tar.gz
Algorithm Hash digest
SHA256 022cea9c88b84d8e5b019ee386e56b3b1f4c20c48b37c302d436583c2954d4c6
MD5 cc5cefdaf08abac1b5b6a38d4e9194ed
BLAKE2b-256 5dcc6f31588fdae60bb5cce89a667c5627d52d0ba1ed49307e86b844fac55e41

See more details on using hashes here.

File details

Details for the file cjm_transcription_core-0.0.7-py3-none-any.whl.

File metadata

File hashes

Hashes for cjm_transcription_core-0.0.7-py3-none-any.whl
Algorithm Hash digest
SHA256 c088c0eb1bf5cca199a12c75ef75e2ebaf7682f726ccbc694d117b8176a2c908
MD5 a90b99e734dfa98ff565998f1e17e2e9
BLAKE2b-256 cbe1461db8c36fcf8fffb44bf04bbb1badb9cdddbf1775e5abad650dec4f51ed

See more details on using hashes here.

Release history Release notifications | RSS feed

0.0.9

2 files

0.0.8

2 files

This release

0.0.7 This release

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page