Typed transcription task-adapter interface — TranscriptionAdapter ABC, TranscriptionResult wire DTO, and transcription persistence helpers.
Project description
cjm-transcription-adapter-interface
Install
pip install cjm_transcription_adapter_interface
Project Structure
nbs/
├── adapter.ipynb # The typed transcription task contract — `TranscriptionAdapter` ABC +
├── core.ipynb # Standardized result DTO for the transcription task — wire-registered
└── storage.ipynb # Standardized SQLite storage for transcription results with content hashing
Total: 3 notebooks
Module Dependencies
graph LR
adapter["adapter<br/>Transcription Adapter"]
core["core<br/>Core Data Structures"]
storage["storage<br/>Transcription Storage"]
adapter --> core
1 cross-module dependencies detected
CLI Reference
No CLI commands found in this project.
Module Overview
Detailed documentation for each module in the project:
Transcription Adapter (adapter.ipynb)
The typed transcription task contract —
TranscriptionAdapterABC +
Import
from cjm_transcription_adapter_interface.adapter import (
TranscriptionToolProtocol,
TranscriptionAdapter
)
Classes
@runtime_checkable
class TranscriptionToolProtocol(Protocol):
"""
PROVISIONAL structural contract for transcription-capable tools.
Mirrors the fused-era surface (task-shaped `execute`); re-derived from
native tool surfaces when the Option C cascade splits tools (stage 8).
"""
def execute(self, audio: Union[str, Path], **kwargs) -> Any: ...
class TranscriptionAdapter(TaskAdapter):
"""
Typed transcription task adapter: model-ready audio in,
`TranscriptionResult` out.
Input contract (carried over from the fused-era TranscriptionPlugin):
the caller guarantees MODEL-READY audio — format / sample-rate /
channel handling happens upstream (ffmpeg `convert_for_model`), never
in the adapter.
Persistence sits BESIDE the task method (pass-2 Thread 3): the storage
module's `TranscriptionStorage` provides the adapter-level cache /
persist seam (`get_cached(audio_path, audio_hash, config_hash)` +
`save_with_logging(...)`).
The result DTO is wire-registered ("transcription.result"): returned
values cross the worker boundary typed via the substrate's `core.wire`
envelope.
"""
def transcribe(
self,
audio: Union[str, Path], # Path to MODEL-READY audio (converted upstream)
**kwargs, # Adapter-specific options (language, task, ...)
) -> TranscriptionResult: # Typed transcription output
"Transcribe model-ready audio to text."
Core Data Structures (core.ipynb)
Standardized result DTO for the transcription task — wire-registered
Import
from cjm_transcription_adapter_interface.core import (
TranscriptionResult
)
Classes
@dataclass
class TranscriptionResult:
"Standardized output for all transcription plugins."
text: str # The transcribed text
confidence: Optional[float] # Overall confidence (0.0 to 1.0)
segments: Optional[List[Dict[str, Any]]] # Timestamped segments
metadata: Dict[str, Any] = field(...) # Additional metadata
Transcription Storage (storage.ipynb)
Standardized SQLite storage for transcription results with content hashing
Import
from cjm_transcription_adapter_interface.storage import (
TranscriptionRow,
TranscriptionStorage
)
Classes
@dataclass
class TranscriptionRow:
"A single row from the transcriptions table."
job_id: str # Unique job identifier
audio_path: str # Path to the source audio file
audio_hash: str # Hash of source audio in "algo:hexdigest" format
config_hash: str # Hash of the effective transcription config used
text: str # Transcribed text output
text_hash: str # Hash of transcribed text in "algo:hexdigest" format
segments: Optional[List[Dict[str, Any]]] # Timestamped segments
metadata: Optional[Dict[str, Any]] # Plugin metadata
created_at: Optional[float] # Unix timestamp
class TranscriptionStorage:
def __init__(
self,
db_path: str # Absolute path to the SQLite database file
)
"Standardized SQLite storage for transcription results."
def __init__(
self,
db_path: str # Absolute path to the SQLite database file
)
"Initialize storage, create table, run migrations, and build indexes."
def save(
self,
job_id: str, # Unique job identifier
audio_path: str, # Path to the source audio file
audio_hash: str, # Hash of source audio in "algo:hexdigest" format
config_hash: str, # Hash of the effective transcription config
text: str, # Transcribed text output
text_hash: str, # Hash of transcribed text in "algo:hexdigest" format
segments: Optional[List[Dict[str, Any]]] = None, # Timestamped segments
metadata: Optional[Dict[str, Any]] = None # Plugin metadata
) -> None
"Save or replace a transcription result (upsert by audio_path + config_hash)."
def save_with_logging(
self,
*,
job_id: str, # Unique job identifier
audio_path: str, # Path to the source audio file
audio_hash: str, # Hash of source audio in "algo:hexdigest" format
config_hash: str, # Hash of the effective transcription config
text: str, # Transcribed text output
text_hash: str, # Hash of transcribed text in "algo:hexdigest" format
segments: Optional[List[Dict[str, Any]]] = None, # Timestamped segments
metadata: Optional[Dict[str, Any]] = None, # Plugin metadata
logger: Optional[logging.Logger] = None # Optional logger for success/failure messages
) -> bool: # True if saved; False if the save failed (error logged, not raised)
"Save a result, logging success/failure. Failures are logged and swallowed (returns False).
Centralizes the try/save/log/except block every transcription plugin reimplements.
Returns True on success so callers can gate post-save side effects on the result."
def get_cached(
self,
audio_path: str, # Path to the source audio file
audio_hash: str, # Content hash of the audio (cache miss if the file changed)
config_hash: str # Hash of the effective transcription config
) -> Optional[TranscriptionRow]: # Cached row or None
"Retrieve a content-correct cached transcription result.
Matches on audio_path + audio_hash + config_hash. A changed audio file
(new audio_hash) misses even if a stale row exists at the same
(audio_path, config_hash) — the next save() replaces it."
def get_by_job_id(
self,
job_id: str # Job identifier to look up
) -> Optional[TranscriptionRow]: # Row or None if not found
"Retrieve a transcription result by job ID."
def list_jobs(
self,
limit: int = 100 # Maximum number of rows to return
) -> List[TranscriptionRow]: # List of transcription rows
"List transcription jobs ordered by creation time (newest first)."
def verify_audio(
self,
job_id: str # Job identifier to verify
) -> Optional[bool]: # True if audio matches, False if tampered, None if job not found
"Verify the source audio file still matches its stored hash."
def verify_text(
self,
job_id: str # Job identifier to verify
) -> Optional[bool]: # True if text matches, False if tampered, None if job not found
"Verify the transcription text still matches its stored hash."
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cjm_transcription_adapter_interface-0.0.2.tar.gz.
File metadata
- Download URL: cjm_transcription_adapter_interface-0.0.2.tar.gz
- Upload date:
- Size: 12.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
195b923d904defa5b28a2c339af9c9033dc19d1d03f56153efc3c34c31ddbe4c
|
|
| MD5 |
e1cc528773e201472af224c292b6cb59
|
|
| BLAKE2b-256 |
e1843b2771f9a4ebae4ef82da930bc775c9b2c3b7a40151c9803eafd493a5bfb
|
File details
Details for the file cjm_transcription_adapter_interface-0.0.2-py3-none-any.whl.
File metadata
- Download URL: cjm_transcription_adapter_interface-0.0.2-py3-none-any.whl
- Upload date:
- Size: 14.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2f48a05e0642fac7fbc4f444955df8a48954765301cfecb9df5c0ede955195ee
|
|
| MD5 |
f0d2656d77bc12799f489f1f52550042
|
|
| BLAKE2b-256 |
f1553b4a2882108f182399f7ecc364a2edccb829108a8ee6e5ba91ebb2f34b8a
|