Skip to main content

textflowkit

Cross-platform media transcription toolkit. One core, one CLI, thin adapters.

Project landing page · PyPI package · GitHub releases · User manual

Current release: v0.1.6.

The static site deployment is hosted on Cloudflare Pages. GitHub remains the source and CI host; the website does not run the transcription engine.

Paste a URL or point at a file; get timestamped transcripts and subtitle files back. Built as a reusable primitive for developers — designed to sit under multiple products, AI harnesses, and agents.


Quickstart

Requires Python ≥ 3.10 and ffmpeg on PATH.

Windows (PowerShell or CMD):

python -m pip install textflowkit
textflowkit doctor
textflowkit transcribe .\meeting.mp4 --formats srt,txt --output-dir .\out

macOS or Linux:

python -m pip install textflowkit
textflowkit doctor
textflowkit transcribe ./meeting.mp4 --formats srt,txt --output-dir ./out

doctor prints the Python, ffmpeg, yt-dlp, JavaScript-runtime, and optional-extra versions this install will use. The transcribe example writes the SRT and TXT files into out\ and prints the path of each one; give it a supported URL instead of a path to fetch remote media. See Install below for the extras, and Usage for language, translation, speaker labels, and resume.


What it does

URL or file  ─►  detect platform  ─►  acquire media  ─►  ffmpeg
                                                          │
                              ┌───────────────────────────┘
                              ▼
                    speech-to-text (Whisper)
                              │
                              ▼
                 canonical transcript (JSON)
                              │
        ┌─────────┬───────────┼───────────┬──────────┐
        ▼         ▼           ▼           ▼          ▼
       TXT       SRT         VTT        JSON     Markdown

DOCX and PDF are available through the optional export extra.

Architecture

The design principle is one engine, three doors. Everything of substance lives in the core; the interfaces are thin.

Layer Path Responsibility
Core src/textflowkit/core Canonical transcript model, pipeline orchestration
Sources src/textflowkit/sources Per-platform URL normalization + media acquisition
Renderers src/textflowkit/render TXT / SRT / VTT / JSON / Markdown / DOCX / PDF output
CLI src/textflowkit/cli.py Reference interface (subprocess-friendly)
MCP src/textflowkit/adapters/mcp_server.py stdio + Streamable HTTP, for AI harnesses
HTTP src/textflowkit/adapters/http_server.py JSON API, for software products and web frontends

Because the core owns the pipeline, adding a door is cheap — and adding a platform means writing one source adapter, not another tool.

Recognized sources

YouTube · TikTok · Facebook · Instagram · Vimeo · Twitch · Bilibili · Rumble · Kick · Zoom · Medal · Loom · Dropbox — plus direct media URLs and local files.

All 13 are recognised through yt-dlp; only YouTube has an opt-in, maintained live end-to-end smoke. It is run on a Windows maintainer machine before a release, not on GitHub-hosted runners or every pull request. Local files have also been transcribed live. The other 12 are not release-verified end to end, and some sources require cookies or change their access rules frequently. See docs/sources.md.

Install

Requires Python ≥ 3.10 and ffmpeg on PATH.

python -m pip install textflowkit
textflowkit doctor

For MCP or the JSON HTTP adapter, install the matching extra:

python -m pip install 'textflowkit[mcp]'    # MCP server
python -m pip install 'textflowkit[http]'   # JSON HTTP API

PDF/DOCX export is optional. The default wheel stays small; the export extra installs textflowkit-fonts for offline multilingual PDF rendering:

python -m pip install 'textflowkit[export]'

Python API

The same pipeline used by the CLI and adapters is available to Python callers:

from textflowkit import transcribe

result = transcribe(
    "meeting.mp4",            # also accepts supported URLs
    model="small",
    formats=["json", "srt", "txt"],
    output_dir="transcripts", # omit to return the transcript without writing files
)
print(result.transcript.text)
print(result.transcript.duration)  # full media duration in seconds
print(result.outputs)              # pathlib.Path objects for written files
for segment in result.transcript.segments:
    print(segment.start, segment.end, segment.speaker, segment.text)
    for word in segment.words:
        print("  ", word.start, word.end, word.text)

transcribe() returns TranscribeResult with a canonical Transcript and written output paths. Transcript.to_dict() / .to_json() preserve segment and word timing; older transcript JSON without words remains readable. Pass input_root= to confine local input paths for untrusted callers. See the install guide for ffmpeg and Windows ROCm setup.

MCP and HTTP transcript reads omit word timings by default to keep responses small; set include_words=true on a JSON read to receive them. Saved files, Python results, and durable job records still retain the source-language words, including when segment text has been translated.

AMD ROCm on native Windows: do not use the generic command in an environment with a working ROCm PyTorch install. Ordinary dependency resolution can replace that torch build. Follow the ROCm install notes to preserve it.

The same version's wheel and source archive are also on the GitHub release page. Starting with the next release, every release published by this project's release workflow carries a SHA256SUMS asset listing the SHA-256 hash of each wheel and source archive it contains, so you can check a download with sha256sum -c SHA256SUMS. Releases published before that change, v0.1.5 included, have no such asset. For editable source development, see CONTRIBUTING.md.

Usage

# transcribe a URL or a local file
textflowkit transcribe "https://www.youtube.com/watch?v=..."

# pick formats and an output directory
textflowkit transcribe ./talk.mp4 --formats srt,vtt,txt,json --output-dir ./out

# force a language instead of auto-detecting
textflowkit transcribe "$URL" --language en

# translate the transcript (uses the configured backend)
textflowkit transcribe "$URL" --translate-to Spanish

# label speakers (requires the optional extra and a Hugging Face token)
textflowkit transcribe "$URL" --diarize

# resume a previous run instead of starting over
textflowkit transcribe "$URL" --resume

Resume reuses completed work from an earlier run. It needs two things: the same source, model, language, and options as the original run, and a durable job store (TEXTFLOWKIT_DB) - a checkpoint cannot outlive a process that kept it in memory. For a local file, a completed-job resume checks the file still exists and matches its checkpointed normalized path, size, and SHA-256 content digest. Missing or changed files fail with an actionable error; v0.1.1-era local checkpoints without a fingerprint must be resubmitted without --resume. For URLs, resume deliberately reuses the saved transcript by URL/options; it does not assert that the remote bytes are still identical.

# re-render an existing transcript in another format
textflowkit export ./transcript.json --format vtt

Use as an MCP server

python -m pip install 'textflowkit[mcp]'
textflowkit-mcp                                  # stdio
textflowkit-mcp --transport http --port 8766     # Streamable HTTP

Tools: transcribe_media, submit_batch_media, resume_job, get_job_status, get_transcript, export_transcript, list_sources, list_jobs, cancel_job, search_transcript.

Connection-smoked against DSH, Claude Code, OpenCode, and Codex desktop. These checks are not end-to-end transcription runs driven by each harness:

  • DSH - the server spawned as a child of the harness's MCP client, which then completed an MCP handshake, discovered the tools, and returned real data from a list_sources call.

  • Claude Code - claude mcp list reports textflowkit: √ Connected (stdio).

  • OpenCode - opencode mcp list reports textflowkit connected over Streamable HTTP.

  • Codex desktop - after repair of an unrelated model-catalog issue, a live list_jobs tool call succeeded. This does not prove every tool or a full transcription in Codex.

There is also a protocol test that launches the server as a real subprocess and speaks newline-delimited JSON-RPC over stdio, so the entry point, framing, and version negotiation are covered on every CI run (tests/test_stdio_protocol.py). See docs/adapters.md for per-harness configuration.

Use as an HTTP API

python -m pip install 'textflowkit[http]'
textflowkit-http --port 8767

Submit a job, poll it, fetch the transcript. Developer mode is unauthenticated and defaults to localhost. The opt-in JSON HTTP production profile requires a Bearer token, explicit roots, durable SQLite jobs, and request/rate/media/output limits; URL input additionally requires an SSRF-filtering egress proxy. See adapter deployment details.

A Dockerfile and a compose.yaml for that profile sit at the repository root: ffmpeg and a JavaScript runtime in the image, a nonroot runtime user, the host port published to loopback only, and a token the operator supplies (there is no default). This is a Linux deployment example, not a published service - it ships no TLS gateway and no egress proxy, so URL jobs fail closed until you provide one. The image has not been built or started anywhere: only static contract checks cover it, and running an actual build is deliberately out of scope here, so no image is claimed to build or start. See the container example.

Durable, bounded, cancellable

TEXTFLOWKIT_DB=./jobs.db            # job state survives restart (SQLite)
TEXTFLOWKIT_MAX_CONCURRENCY=1        # default; Whisper saturates a GPU alone

cancel_job stops a queued job immediately, or a running job at its next stage boundary. See docs/adapters.md.

Long jobs never block

The MCP and HTTP adapters are job-based: submission returns a job id immediately and clients poll for completion. The CLI uses the same core but waits for the result. This lets AI harnesses, software products, and a future web frontend share the job contract without blocking a request.

Status

v0.1.6 release. Core, CLI, MCP, and HTTP have automated coverage. This release is the post-v0.1.5 review repair set: security hardening for media acquisition and the HTTP and MCP adapters, safer subtitle wrapping and output-file publication, job and checkpoint storage corrections, an opt-in faster-whisper engine, a container example, and fail-closed release guards for tagged versions, README claims, and exact-commit CI. The v0.1.5 release added a speech-bearing self-test, full-media duration, retained word timings, optional PDF fonts, and tokenless PyPI publishing, and its live evidence is the only live evidence recorded on this page. Windows-native ROCm and a local Windows YouTube run on the v0.1.5 release commit were verified. That run's receipt records the clip, timestamps, and hashes; the checklist's recognized-speech assertion was added after it, so treat it as shape evidence for that commit rather than proof of what was said. The GitHub-hosted YouTube attempt was blocked by a bot challenge, so hosted live transcription is not verified. No v0.1.6 tag, upload, or live run is verified here. This does not imply that all 13 platforms or every harness workflow has been tested end to end. See the user manual and roadmap.


⚠️ NO WARRANTY — AS IS

This software is provided "AS IS", WITHOUT WARRANTY OF ANY KIND, express or implied, including but not limited to the warranties of MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, and NONINFRINGEMENT. See LICENSE (Apache-2.0, §7–8) for the full disclaimer and limitation of liability.

You are responsible for what you transcribe. textflowkit can fetch media from third-party platforms. Copyright, terms-of-service, and privacy obligations for any media you choose to process are yours alone. See LEGAL.md.

License

Apache-2.0 — see LICENSE. Includes an explicit patent grant and a limitation of liability.

Contributing

Issues and PRs welcome. Please read LEGAL.md before adding a source adapter.

Release files for textflowkit 0.1.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for textflowkit 0.1.6
File Size Uploaded
textflowkit-0.1.6.tar.gz 488.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for textflowkit 0.1.6
File Interpreter ABI Platform
textflowkit-0.1.6-py3-none-any.whl Python 3 none any Details

Total release size: 727.2 kB

Release files / textflowkit-0.1.6.tar.gz

Download URL textflowkit-0.1.6.tar.gz
Size 488.6 kB
Tags Source
SHA-256 checksum
How to use checksums
eab81c1800b0c37b0ecfdea339d075331f5d27712c64c2752ec918e3fe1b3e92
BLAKE2b-256 checksum
How to use checksums
c570f10f5101f3d72b5a1ccd9ae13610cfb33e0aa81e75eb8b918d39820b8821
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log

Release files / textflowkit-0.1.6-py3-none-any.whl

Download URL textflowkit-0.1.6-py3-none-any.whl
Size 238.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1d0db0864e49dd82a15407e4aedb4c27e1754747ffb10e7eb75b65fec2836b1a
BLAKE2b-256 checksum
How to use checksums
a8b8122432b5d1df7970ea4fd3e6d0d3243301a27c046b4e531d3ce5fa447128
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.6 This release

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page