Skip to main content

yt-notes

A Python CLI that saves YouTube captions, readable research notes, and structured JSON with evidence linking each extracted item to the original video's timestamps. It accepts individual videos, channels, and playlists without a YouTube API key.

Install

From a cloned or unpacked checkout on glibc Linux or macOS 13+, x64 or ARM64:

bash install.sh
yt-notes
yt-notes doctor

The script installs the CLI and terminal UI together. It downloads private uv, Python 3.12, and Bun 1.4.2 build tools, compiles the frontend, and installs matching wheels from this checkout. You do not need Python, Bun, Node.js, an activated virtualenv, GitHub Actions artifacts, or an existing dist/ directory. Internet access, Bash, curl, tar, and normal system utilities are required; sudo is not used. Alpine/musl and native Windows are not supported by the shell installer.

The launcher is ~/.local/bin/yt-notes. The installer adds its directory to your Bash/Zsh startup files when needed and prints the command for your current shell. An activated virtualenv can still select its old yt-notes: run deactivate, then hash -r, or use the absolute launcher printed by the installer.

bash install.sh --cli-only    # CLI without the terminal UI or Bun build
bash install.sh --asr         # Include faster-whisper; no model weights downloaded

Re-run the script after updating your checkout to upgrade or repair the install, including changes that keep the same package version. Every run builds fresh wheels and verifies the installed packages before switching the launcher. Failures leave the previous installation usable. Installed environments remain independent of the checkout, so moving it does not break the command.

Runtime files are stored in ${XDG_DATA_HOME:-$HOME/.local/share}/yt-notes/install. Set YT_NOTES_INSTALL_DIR and YT_NOTES_BIN_DIR to override those locations. Previous environments remain in install/envs/ so an already-running session can finish; the install/current symlink identifies the active one. Configuration, credentials, history, caches, and outputs keep their existing locations.

Run yt-notes doctor to inspect optional runtimes and credential presence. Deno, FFmpeg, Ollama, API keys, and model weights are configured separately when needed.

Manual Python installation and development

For a CLI-only install with Python 3.11 or newer, or on Windows:

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install .
yt-notes --help
yt-notes doctor

On Windows, activate with .venv\Scripts\Activate.ps1. For editable development, use python -m pip install -e '.[dev]'. The CLI does not install software or pull local models automatically. Some YouTube extraction paths need a supported JavaScript runtime; install Deno when yt-dlp reports it is required. The package includes yt-dlp's EJS dependency.

Terminal workspace

The optional interface uses OpenTUI + SolidJS + TypeScript, the stack used by OpenCode. Python still owns configuration, discovery, providers, caching, and exports. The version-matched yt-notes-ui companion wheel bundles the frontend, runtime, and native renderer. Users do not need Bun or Node.js.

The recommended bash install.sh flow builds and installs the UI automatically. For a manual wheel install, download and unpack the yt-notes-<platform> artifact from the repository's UI wheels GitHub Actions run into dist/, then install both wheels:

python -m pip install --find-links dist 'yt-notes[tui]==0.2.1'
yt-notes                            # opens the workspace in a terminal
yt-notes tui                        # explicit alias
yt-notes --config config.example.toml --provider ollama tui

--find-links dist only searches wheels that already exist in that directory; it does not download CI artifacts or build the missing companion package. If you see No matching distribution found for yt-notes-ui, use bash install.sh from your checkout, or download and unpack both matching wheels first.

CI builds Linux glibc x64/arm64, macOS x64/arm64, and Windows x64 wheels. Linux artifacts are built on Ubuntu 22.04; macOS needs 13 or newer. Alpine/musl is not supported. Wheels are CI artifacts; publishing to PyPI is a separate release action. Once published, python -m pip install 'yt-notes[tui]' also works without --find-links. Launching never downloads dependencies. A missing or mismatched UI package produces installation instructions.

Use the workspace

  1. Paste a YouTube video, channel, or playlist URL and press Enter.
  2. For collections, use arrows and Space to choose videos. Tab focuses search; Ctrl+A selects the filtered list. Press Enter to review.
  3. Review the model, language, formats, and destination; choose Start run. /settings also includes transcript-only mode, ASR and its model, caption languages, collection limit, refresh, overwrite, and provider endpoint.
  4. Watch the activity feed, per-video cost, and video queue. The highlighted video's thumbnail appears in supported terminal layouts; unavailable artwork falls back to text. After completion, Tab to a generated file and press Enter for a scrollable preview. The batch manifest path appears in the feed; generated paths can be copied from the terminal.
  5. The home screen shows the latest three runs and their API costs. Open /history to browse all available runs and inspect each outcome. Press R in a run's details to review its failed videos before retrying.
Command / key Action
/model Choose provider and enter any supported model ID
/settings Edit run options for this session
/output Set the destination for notes and batch manifests
/history Browse past runs and inspect outcomes
/retry Review failures from the most recent run with failures
/import Enter a saved yt-notes video JSON export path
/new Return home for another source
/help, /quit Show shortcuts or exit safely
Ctrl+P Search all commands from any screen
Up / Down, Enter Choose and confirm
Tab / Shift+Tab Change focus; Tab completes palette commands
Esc Close an overlay or leave review/selection
Ctrl+B Toggle video queue; it collapses on narrow terminals
PgUp / PgDn, Up / Down Scroll the focused feed or preview
Ctrl+C Request cancellation during work; quit when idle

Slash commands work at the home prompt; Ctrl+P works throughout the workspace. /output accepts absolute paths, paths relative to the launch directory, and ~ for your home directory. Use /settings → Save defaults to keep the destination for future sessions. CLI runs can use --output-dir PATH. Ordinary letters, including q, remain normal input. The layout adapts to light/dark terminal preferences and resizing. Below 50×16, a resize notice appears while active processing continues.

Cancellation, imports, and recovery

Cancellation finishes the current video and skips the remaining queue. The banner stays visible until Python finishes saving. The manifest records interrupted and skipped_video_ids when videos remain; stopping during the final video can still complete the batch. /quit during work waits for this same safe stopping point. A disconnected frontend also requests a graceful stop. Network retries or ASR can make the current video take time; force-killing Python cannot guarantee its final checkpoint.

/import accepts a per-video JSON export, including transcript checkpoints from failed summaries, and follows the same review flow. It does not contact YouTube; refresh must be off. A batch manifest is an index of outcomes, not an importable transcript. Completed files remain available after cancellation; a later run can reuse existing outputs and cached work.

Every new CLI or TUI batch is indexed locally, so history remains available after changing output directories. The current output directory is also scanned for older batch manifests. Older manifests can be browsed, but retries use the current settings because those manifests lack a settings snapshot. A retry starts a new batch containing only failed videos, with the original settings and destination prefilled when available; credentials are read from current configuration and are never stored in history. In the shell, use yt-notes retry for the latest run with failures, or yt-notes retry --run /path/to/batch_MANIFEST.json for a specific manifest. The command lists failures and asks for confirmation; add --yes in scripts. --output-dir PATH overrides the saved destination for that retry.

Each batch writes a complete JSON-lines diagnostic log beside its manifest. The terminal workspace also writes a session log in the local yt-notes data directory, covering discovery and setup errors before a batch starts. Logs contain stages, retry attempts, response status codes, usage and safe error types; they omit credentials, transcripts, prompts, signed media URLs and raw provider responses. Copy the newest log on a Wayland desktop with:

yt-notes debug-log --latest | wl-copy

Use yt-notes debug-log --run /path/to/batch_MANIFEST.json to print a specific batch's complete log. Omit | wl-copy when you want to inspect the text first.

API keys, cookies, and proxies are represented only by presence indicators. Set credentials through the environment, .env, or existing configuration mechanism. Select /settings → Provider API keys, choose OpenRouter, paste the key without exposing it on screen, and press Enter. Press Enter again to save, or return to settings and choose Save defaults and entered keys to .env. This creates or updates .env in the directory where you launched yt-notes, preserves unrelated lines, and restricts the file to your account on POSIX systems. Blank key fields leave saved keys unchanged. .env contains plain text secrets: keep it private and out of source control. A key entered in the UI applies immediately to this session; exported environment variables and an explicit config file keep precedence on later launches. Provider response bodies and library diagnostics are never rendered. Options passed before tui override settings for that session.

The regular CLI works without the optional package. --url and --from-json retain their behavior; explicit processing options without a source retain the existing interactive URL prompt. With no TTY, provide an explicit source.

Quickstart

# Open the workspace (requires the optional UI wheel).
export OPENAI_API_KEY="your-key"
yt-notes

# Summarize one video, or save captions without calling any LLM.
yt-notes --url "https://www.youtube.com/watch?v=VIDEO_ID"
yt-notes --url "https://www.youtube.com/watch?v=VIDEO_ID" --transcript-only

# Latest 20 regular uploads; Space selects, Enter confirms.
yt-notes --url "https://www.youtube.com/@CHANNEL"

# Enumerate all regular uploads before selecting; this can take time.
yt-notes --url "https://www.youtube.com/@CHANNEL" --limit 0

# Noninteractive collection selection, in source order.
yt-notes --url "https://www.youtube.com/@CHANNEL" --limit 50 --select "1,3,5-8"
yt-notes --url "https://www.youtube.com/playlist?list=PLAYLIST_ID" --all

# Choose caption preferences, summary language, and formats.
yt-notes --url "VIDEO_URL" --language en,hi --summary-language English --format md,json,txt

Replace the placeholder URLs/IDs with real ones. --all selects the retrieved entries; add --limit 0 to retrieve the entire collection. Channel roots use /videos; explicitly supplied /shorts and /streams tabs are respected. Playlist order is preserved and duplicate IDs are removed. A video URL with &list=... still processes only that video. Active/upcoming streams and archives still processing are reported as unavailable for this run.

Listing enrichment is bounded (20 entries by default); unavailable dates and durations display as unknown. Selected videos receive full metadata retrieval. Without a TTY, provide --url/--from-json and --all or --select for collections.

LLM providers

Provider Default model Credentials / setup
OpenAI (default) gpt-5.4-mini OPENAI_API_KEY
OpenRouter openai/gpt-4.1-mini OPENROUTER_API_KEY
Anthropic claude-sonnet-4-6 ANTHROPIC_API_KEY
Ollama qwen3:8b Running local Ollama and an installed model
export OPENROUTER_API_KEY="your-key"
yt-notes --url "VIDEO_URL" --provider openrouter --model openai/gpt-4.1-mini
yt-notes --provider openrouter doctor

export ANTHROPIC_API_KEY="your-key"
yt-notes --url "VIDEO_URL" --provider anthropic

# Install Ollama separately, then explicitly download the model.
ollama pull qwen3:8b
yt-notes --url "VIDEO_URL" --provider ollama
yt-notes --provider ollama doctor

# Models and endpoint roots can be overridden.
yt-notes --url "VIDEO_URL" --provider ollama --model YOUR_MODEL \
  --base-url http://localhost:11434

For OpenRouter, pass a model ID from the model catalog as --model author/model. The model and its serving endpoint must support JSON schema structured outputs. The CLI sends require_parameters: true so OpenRouter routes to compatible endpoints. You can select any compatible model ID; no hardcoded model list limits selection. Unsupported models, insufficient credits, or key spending limits produce an error. The chosen model is saved in the output and cache identity.

To keep your selection between runs, set YT_NOTES_PROVIDER=openrouter and YT_NOTES_MODEL=author/model, or put provider = "openrouter" and model = "author/model" in your explicitly supplied TOML config. --model takes precedence. OpenRouter uses https://openrouter.ai/api/v1 by default; allow openrouter.ai through any network allowlist. No additional AI SDK is needed: the Python CLI calls the documented HTTP API using its existing HTTPX transport.

Cloud providers receive the transcript content needed for extraction. Ollama keeps LLM inference local; YouTube retrieval still requires a network connection. The tool never switches providers automatically. OpenAI requests use store=false. Ollama runs with explicit context/output budgets and thinking disabled; use a model that supports structured JSON. Larger context windows increase local memory use.

API keys are checked before summary processing; transcript-only mode needs no LLM credentials. Invalid model access, refusal, output truncation, or repeated schema failure is reported rather than written as a successful summary. Token usage is provider-reported. OpenRouter charges use its reported usage.cost. Other cloud charges are labelled estimates based on token usage and a saved price snapshot; missing usage or pricing is shown as partial or unavailable. The default rate table covers gpt-5.4-mini and claude-sonnet-4-6 at their published OpenAI and Anthropic rates as of 2026-10-08. For other models or custom endpoints, set YT_NOTES_PRICE_INPUT_PER_MILLION and YT_NOTES_PRICE_OUTPUT_PER_MILLION (USD), plus YT_NOTES_PRICE_PROVIDER and YT_NOTES_PRICE_MODEL to bind those rates to the exact model. The matching price_* fields work in TOML; the workspace binds rates automatically when edited there. Optional cached-input and cache-write rates use the same naming pattern. Ollama and transcript-only runs show $0 in API charges; local compute and ASR hardware costs are not included. Cached notes add no new API charge.

Outputs and evidence

The default directory is ./outputs, relative to the invocation directory:

outputs/
  YYYY-MM-DD_VIDEO_ID.md
  YYYY-MM-DD_VIDEO_ID.json
  YYYY-MM-DD_VIDEO_ID.txt       # optional --format txt,md,json
  batch_TIMESTAMP_RANDOM.json
  batch_TIMESTAMP_RANDOM.debug.jsonl  # sanitized per-run diagnostics
  .state/                     # canonical records and per-video process locks

Batch manifest version 1.2 includes each attempt's usage and cost alongside the selected video details and nonsecret run settings needed for retry. Versions 1.0 and 1.1 remain readable; their costs are unavailable unless recorded.

Markdown includes YAML frontmatter, executive summary, takeaways, chronological topics, explicit actions, speaker claims, tools/resources, and the full transcript. Every extracted item has timestamp links computed from its cited segment IDs.

JSON schema version 1.0 contains:

  • video: curated metadata, including upload date when known.
  • A validated channel avatar URL when YouTube supplies one; video thumbnails are derived from the public video ID for the terminal preview.
  • transcript: language, backend, manual/generated status, original cues and separate normalized cues. Original text/timestamps are retained.
  • chunks: bounded inputs with text, source IDs and time spans for downstream RAG.
  • extraction: executive summary, takeaways, topics, actions, claims and resources.
  • evidence: source cue ID → start/end seconds and YouTube playback URL.
  • Processing date, transcript hash, configuration fingerprint, prompt/model identifiers, usage, warnings and errors.

Statuses are pending, transcript_only, complete, and failed. Pending files are checkpoints, not completed notes. Missing values use nulls/empty arrays. Batch manifests record success, cached reuse or failure for every attempted video. They also record interruption and artifact paths. The Markdown date is the processing date, separate from YouTube's upload date; reused outputs retain their original processing date.

See examples/ for explicitly fictional sample data and outputs.

Recovery, reprocessing and configuration

# Repeat the same command to reuse source and successful LLM stages.
yt-notes --url "VIDEO_URL"

# Refetch YouTube content. Existing completed reports are protected.
yt-notes --url "VIDEO_URL" --refresh --overwrite

# Regenerate entirely from a saved transcript, without contacting YouTube.
yt-notes --from-json "outputs/YYYY-MM-DD_VIDEO_ID.json" \
  --provider anthropic --output-dir alternative-notes

# Explicit configuration file.
yt-notes --config config.example.toml --url "VIDEO_URL"

Precedence is CLI flags → YT_NOTES_* environment variables → an explicitly supplied TOML file → .env → defaults. For example, YT_NOTES_LANGUAGES=en,hi, YT_NOTES_TRANSCRIPT_ONLY=true, YT_NOTES_FORMATS=md,json, and YT_NOTES_CACHE_DIR=/path/to/cache. Standard OPENAI_API_KEY, OPENROUTER_API_KEY, and ANTHROPIC_API_KEY take precedence over their prefixed equivalents. Matching variables can also be stored in a project .env file; exported environment variables override .env. OLLAMA_HOST is used for Ollama if no base URL is configured. Keep .env private and out of version control. Keep model/base_url unset in provider-neutral configs so changing provider picks the matching defaults. Unknown TOML fields fail validation.

Raw transcripts and validated intermediate extractions are cached under the platform's user cache directory (~/.cache/yt-notes on typical Linux systems). Cache identity includes source preferences, transcript content, model/endpoint, prompt version and relevant processing settings. Corrupt cache entries are misses. Cookies, proxy credentials and API keys are excluded from exported/cache objects. --from-json can reconstruct notes even if the cache was deleted.

Completed reports with different settings/content require --overwrite or a new output directory. Pending/failed runs can resume. A failed refresh preserves an existing completed report; the new failure appears in the batch manifest. Each artifact is atomically replaced and the canonical recovery record is committed last. A file set is not a filesystem-wide transaction: after an abrupt machine shutdown, rerun the same command to reconcile missing/incomplete artifacts.

Captions, language selection and optional ASR

Retrieval tries youtube-transcript-api, then yt-dlp JSON3/WebVTT subtitles. For each preferred language, manual captions precede automatic ones. If no preferred track exists, the video language is tried, then an available original track. Region variants match their base language. No implicit caption translation is requested; --summary-language controls only generated notes.

Cleaning removes subtitle markup and duplicate text from overlapping rolling cues. Repeated speech in distinct non-overlapping cues remains intact. Raw cues are always kept separately.

# Optional local speech recognition when caption retrieval fails.
bash install.sh --asr
yt-notes --url "VIDEO_URL" --asr-fallback
yt-notes --url "VIDEO_URL" --asr-fallback --asr-model medium

For a manually managed Python install, use python -m pip install '.[asr]' instead. ASR defaults to faster-whisper small, CPU/int8. Enabling the flag may download audio and model weights; weights stay in the model library's cache. Audio is temporary and removed after the attempt. No audio is downloaded without the flag. The audio limit is 2 GiB; huge downloads or missing runtime dependencies are reported as failures. ASR results are explicitly marked as generated speech recognition, with the detected language. Install FFmpeg if your selected media format/extraction path requires it. ASR is an optional extra with its own platform wheel requirements; use a supported Python environment if its install fails.

YouTube's unofficial caption endpoints can change or block requests. Metadata and subtitle fallbacks share that platform constraint; they cannot guarantee access. Use --cookies /path/to/cookies.txt (Netscape format, yt-dlp only), or configure an HTTP(S) proxy through YT_NOTES_PROXY. Never commit credential files. Errors and third-party logs are sanitized. Retries apply only to transient failures, with bounded backoff; private/restricted videos and persistent blocks remain failures.

Long videos and extraction quality

All normalized transcript text is covered by bounded chunks. Oversized cues split without losing text, retaining their original source ID. chunk_chars and synthesis_chars count UTF-8 bytes, a conservative multilingual budget, not model tokens. There is no silent transcript truncation. Intermediate notes are cached and recursively combined until final synthesis fits its input budget.

The schema is enforced locally even when the provider supports native structured outputs. One repair attempt is allowed for malformed output, invalid citations, or oversized intermediate notes. Evidence IDs must come from the current chunk or its intermediate inputs. Resource URLs must be present in their cited cues. Actions must be explicit recommendations; absent categories stay empty. Summaries remain model interpretations: valid evidence references establish traceability, not an independent fact check. The full source remains available to inspect.

Jev, Laya and Clef integrations are intentionally deferred; v1 uses generative models for open-ended notes. Optional future passage classification should be evaluated against labeled transcripts before it is allowed to filter content. V1 does not analyze frames/slides, fetch comments, build embeddings, merge videos into a cross-video synthesis, or run a web service.

Development and checks

python -m pip install -e '.[dev]'
pytest
ruff check .
ruff format --check .
mypy
python -m build

Build and develop the frontend

Bun 1.4.2 is a development/build dependency. From the repo:

cd ui
bun install --frozen-lockfile
bun run typecheck
bun test
bun run build
cd ..
python -m build --wheel
python -m build --wheel ui/wheel --outdir dist
python -m pip install --find-links dist -e '.[tui,dev]'
python ui/scripts/smoke_wheel.py
yt-notes

For frontend iteration without rebuilding, set YT_NOTES_UI_PYTHON to the absolute path of the Python interpreter with this project installed, then run bun run dev from ui/. That directory becomes the invocation directory for config, imports, and outputs. Rebuild and reinstall the companion wheel before checking the normal yt-notes launcher.

The build follows OpenTUI's standalone executable procedure. It uses the host platform by default; UI_TARGET and YT_NOTES_UI_WHEEL_TAG allow coordinated cross-build overrides. CI builds on native runners, runs frontend tests and type checks, installs the actual wheels, and exercises the renderer and Python handshake with PATH empty. No workflow publishes packages. Keep the version in the Python project, yt_notes.__version__, UI protocol, UI package, companion wheel, and CI install command synchronized when releasing.

The private bridge uses versioned JSON lines over subprocess stdin/stdout; there is no HTTP server. The Python launcher selects its own interpreter and preserves cwd, environment, explicit config, and CLI overrides. The bridge allows one worker at a time and reuses run_batch for persistence. Frontend tests render and interact with OpenTUI at 120×36, 80×24, and 60×20. Bridge tests cover imports, cancellation, partial failures, precedence, redaction, and disconnect recovery.

The test suite uses deterministic fixtures and mocked networking. It exercises provider wire contracts and evidence validation; it does not claim to benchmark live model accuracy. Review VALIDATION.md for the actual checks performed for this delivery. Live checks are opt-in and are not part of ordinary CI.

requirements-tested.txt records exact dependency versions from the tested Linux Python 3.14 development environment, excluding the optional ASR stack. Use it when reproducing that environment; pyproject.toml is the normal installation contract.

Noninteractive CLI exit codes: 0 success/clean cancellation, 1 processing failure (including partial batches), 2 invalid input/configuration, 130 keyboard interruption. The Python boundaries VideoDiscovery, TranscriptSource, StructuredProvider and StageCache allow replacement services without changing the CLI or schema.

References

Metadata

Release files for yt-notes 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for yt-notes 0.2.1
File Size Uploaded
yt_notes-0.2.1.tar.gz 129.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for yt-notes 0.2.1
File Interpreter ABI Platform
yt_notes-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 198.0 kB

Release files / yt_notes-0.2.1.tar.gz

Download URL yt_notes-0.2.1.tar.gz
Size 129.7 kB
Tags Source
SHA-256 checksum
How to use checksums
f8749cd169646cd2464decadde2ae4508fb3788383cb8ef44412671c24b8c4c4
BLAKE2b-256 checksum
How to use checksums
9cb6a7a98df53525ba2931395efdc853ca74e9d6fe4711e5164d2e8e27d077d1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / yt_notes-0.2.1-py3-none-any.whl

Download URL yt_notes-0.2.1-py3-none-any.whl
Size 68.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
72b38e5194636b093ca179fd7dc5ff48bc442efa835da2c2cdfd206e6d5686f6
BLAKE2b-256 checksum
How to use checksums
8c92d53c31b125fc29130190fdaf8ddc49c2fb02bb8a26914153175260148cb8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page