Podcast Helper
Podcast Helper belongs to a collection of libraries called AI Helpers developed for building Artificial Intelligence.
The Promise
Local-first by design. podcast-helper runs entirely on your machine: it fetches only the episodes and feeds you ask for and processes them locally. Your data is never uploaded to a third-party service, there is no telemetry, no account, no cloud lock-in. Part of the AI Helpers suite: sovereignty over your data through local-first open source.
Universal audio stream consumer for podcasts and any audio-bearing URL. URL in, PCM out (PCM, pulse-code modulation, is the raw stream of uncompressed audio samples that speech-recognition and audio-analysis models expect as input, as opposed to a compressed file like an MP3), for local files, direct audio URLs (RSS enclosure MP3, M4A, Opus, WAV, HLS m3u8), RSS/Atom feed URLs (auto-picks the latest episode), and every yt-dlp-supported source (YouTube, Vimeo, SoundCloud, Twitch VOD or live, and more). Spotify's DRM-gated catalog and Apple Podcasts catalog URLs are refused upfront, with a clear hint toward the RSS-feed workaround.
Documentation
Battle-tested
podcast-helper ships on PyPI under semantic-versioned tags (currently v1.1.5), so a pinned range such as podcast-helper>=1.1.0,<2 stays stable under you. Continuous integration (CI, an automated pipeline that runs the test suite and the linter on every push) is green on main: pytest and ruff check both have to pass before anything merges. It depends on two siblings from the same AI Helpers suite, youtube-helper for the yt-dlp resolution step and os-helper for shared logging and filesystem utilities, and it is itself a dependency of vocal-helper (speaker-labelled transcription), which builds its live and offline audio ingest directly on extract_audio_stream.
Why it exists
Podcast pipelines usually start with the same question, whether the downstream step is automatic speech recognition (ASR), diarization (labeling which speaker is talking when), voice activity detection (VAD, deciding which stretches of audio have speech at all before running the heavier models on them), summarization, or search indexing: "give me a stream of PCM frames from this URL, never mind whether it's a .mp3 link, a feed, a YouTube video, or a podcast hosted on a CDN I've never heard of." This library is that one function, plus the small extras around it (feed, latest_episode) that make working with RSS sources friendly.
Installation
Prerequisites: Python 3.10-3.13, git, and ffmpeg, cross-platform:
- 🍎 macOS (Homebrew):
brew install python git ffmpeg - 🐧 Ubuntu/Debian:
sudo apt update && sudo apt install -y python3 python3-pip git ffmpeg - 🪟 Windows (PowerShell):
winget install Python.Python.3.12 Git.Git Gyan.FFmpeg
We recommend using Python environments. Check this link if you're unfamiliar with setting one up: 🥸 Tech tips.
From PyPI (recommended)
# Core library
pip install podcast-helper
# Optional surfaces
pip install "podcast-helper[cli]"
pip install "podcast-helper[api]"
pip install "podcast-helper[mcp]"
This pulls in youtube-helper (and transitively yt-dlp, os-helper, audio-helper, video-helper), plus feedparser and podcastparser for RSS.
From source (no PyPI)
git clone https://github.com/warith-harchaoui/podcast-helper.git
cd podcast-helper
pip install -e .
# Optional surfaces
pip install -e ".[cli]"
pip install -e ".[api]"
pip install -e ".[mcp]"
Quick start
import asyncio
import podcast_helper as ph
async def main():
# Pass *any* URL: file, direct mp3, RSS feed, YouTube, SoundCloud, Twitch VOD.
async for frame in ph.extract_audio_stream(
"https://feeds.npr.org/510289/podcast.xml", # ← RSS, auto-pick latest episode
target_sample_rate=16000,
to_mono=True,
frame_ms=20,
):
# frame["pcm"]: np.float32 (320,) for 20ms @ 16kHz
# frame["t_abs_s"]: 0.0, 0.02, 0.04, ...
await asr.feed(frame["pcm"])
asyncio.run(main())
For the full catalog of recipes (RSS, yt-dlp sources, live streams, stereo / multichannel, anti-aliasing, downstream ASR / VAD / summarization pipelines), see 📋 EXAMPLES.md.
What URLs are accepted
| Source | Detection | What happens |
|---|---|---|
Local file / file:// |
path exists on disk OR file:// scheme |
ffmpeg opens it directly. |
Direct audio URL (.mp3, .m4a, .opus, .wav, .m3u8, …) |
URL extension is a known audio container | ffmpeg opens it directly with your headers= if any. |
RSS / Atom feed (.xml, .rss, .atom) |
URL extension is a known feed container | podcastparser (fallback: feedparser) parses it; latest episode's enclosure is fetched. |
| YouTube / Vimeo / SoundCloud / Twitch VOD / Twitch live / … | yt-dlp's extractor identifies it | yt-dlp picks bestaudio*, hands the direct URL + headers to ffmpeg. |
| Generic web URL (anything else) | yt-dlp's generic extractor |
URL used as-is. |
| Spotify (open.spotify.com) | hostname match | NotImplementedError: Spotify audio is DRM-gated. Use the show's RSS feed if it exists. |
| Apple Podcasts (podcasts.apple.com) | hostname match | NotImplementedError: Apple URLs point to the catalog, not the audio. Use the show's RSS feed (linked on the show's site, or via getrssfeed.com / Podcast Index). |
Signal-processing correctness
When target_sample_rate differs from the source rate, the conversion is performed by ffmpeg's libswresample (default) or libsoxr (resample_quality="high"). Both apply an anti-aliasing low-pass filter at the new Nyquist frequency (target_sample_rate / 2, the sampling theorem's ceiling on the highest frequency a given rate can represent) before decimation: any content above that ceiling gets removed first, because if it were kept it would fold back down into the audible range as spurious noise once the rate drops, an artifact called aliasing. Naïve subsampling, dropping samples without this filter, skips that step and lets the aliasing through; it is never used here.
Channel handling has exactly two modes; there is no synthetic upmix:
to_mono |
Output shape | What ffmpeg does |
|---|---|---|
True (default) |
(n_samples,) |
Standard downmix (stereo → L+R with -3 dB, 5.1 → ITU mix) |
False |
(n_samples, n_channels) interleaved |
Preserves the source's native channel count |
Working with RSS feeds explicitly
If you want to inspect or select episodes yourself:
import podcast_helper as ph
# Full episode list, most-recent first
episodes = ph.feed("https://feeds.npr.org/510289/podcast.xml", max_episodes=20)
for ep in episodes:
print(ep["published_at"], "|", ep["title"], "|", ep["duration_seconds"], "s")
# Or just the latest one
ep = ph.latest_episode("https://feeds.npr.org/510289/podcast.xml")
print(ep["title"], "→", ep["enclosure_url"])
# Then stream its audio
import asyncio
async def main():
async for frame in ph.extract_audio_stream(ep["enclosure_url"]):
...
asyncio.run(main())
Each Episode dict has a normalized schema regardless of feed flavor:
{guid, title, description, link, published_at (ISO UTC),
duration_seconds, enclosure_url, enclosure_type, enclosure_size_bytes,
image_url}
Multi-surface exposure
podcast-helper exposes the same public functions through five
interchangeable surfaces; pick the one that fits the caller.
| Surface | Entry point | Extra | Best for |
|---|---|---|---|
| Library (async iterator) | import podcast_helper as ph |
none | Python code, notebooks, downstream ASR / VAD / summarization |
| argparse CLI | podcast-helper |
none (stdlib only) | shell scripts, CI, ffmpeg pipelines |
| click CLI | podcast-helper-click |
[cli] |
click-native shells (bash / zsh completion, colored help) |
| FastAPI HTTP | uvicorn podcast_helper.api:app |
[api] |
HTTP microservices, cross-language callers |
| Browser GUI | GET /gui (served by the API) |
[api] |
drop-a-URL episode browser: list · preview · archive, no terminal |
| MCP | podcast-helper-mcp |
[mcp] |
any MCP-aware agent host (calls the same endpoints as tools) |
Install any combination of extras:
pip install 'podcast-helper[cli]' # + click twin
pip install 'podcast-helper[api]' # + FastAPI HTTP surface
pip install 'podcast-helper[mcp]' # + MCP tools (needs [api] too, pulled in automatically)
pip install 'podcast-helper[cli,api,mcp]' # everything
Every surface publishes the same verbs, feed, latest, stream,
record, probe, with identical argument names, so switching
between them is a copy-paste. The Dockerfile in this repo ships the
FastAPI surface on port 8000 out of the box (docker build -t podcast-helper . && docker run --rm -p 8000:8000 podcast-helper).
Browser GUI: the episode browser (GET /gui)
With the [api] extra, the FastAPI app serves a self-contained
single-page episode browser (Tailwind via CDN + vanilla JS, no
build step) that drives the very same endpoints:
pip install 'podcast-helper[api]'
uvicorn podcast_helper.api:app --port 8000
# open http://localhost:8000/gui (or just http://localhost:8000/)
Paste a feed / RSS / audio / yt-dlp URL → List episodes (calls
/feed) → click an episode to see its metadata and play the enclosure
inline → Record to file (calls /record) to download a compressed
archive. Probe classifies any URL. Nothing is uploaded: playback
streams the enclosure straight to your browser, and archiving runs
ffmpeg on your own machine.
For the exhaustive catalog of triggers, phrasings, accepted URLs and
when not to reach for podcast-helper, see
TRIGGERS.md.
For an ambitious visual product on top, see GUI.md.
Live streams
For YouTube / Twitch live URLs, the resolved direct URL is typically an HLS (HTTP Live Streaming, a protocol that serves audio or video as a sequence of small segment files listed in a .m3u8 playlist) manifest. extract_audio_stream detects this (is_live=True) and automatically disables -re real-time pacing (the source paces itself). The async iterator runs indefinitely until the live stream ends; callers should break when they're done.
speed != 1.0 for live streams raises ValueError: you can't fast-forward beyond the live edge. Use speed=... on VOD only.
Not built yet
start_instant/end_instantfor VOD seek.apple_podcasts_to_rss(url), resolving an Apple Podcasts catalog URL to its RSS feed via the iTunes Search API.- Podcast Index API integration.
- Chapters (ID3 CTOC/CHAP, Podcasting 2.0
<podcast:chapters>) and transcripts. - OPML import/export.
See CHANGELOG.md for what has already shipped.
Author
Acknowledgements
Special thanks to Mohamed Chelali and Bachir Zerroug for fruitful discussions.
License
This project is licensed under the BSD-3-Clause License. See the LICENSE file for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file podcast_helper-1.1.5.tar.gz.
File metadata
- Download URL: podcast_helper-1.1.5.tar.gz
- Upload date:
- Size: 57.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ca357881e3df5c7f96f77353927600ff86e3108e792fd6cc183b8900203d0910
|
|
| MD5 |
c0204de49b64e61fce12e46708c677e2
|
|
| BLAKE2b-256 |
49cd7f1fb7c7fdfe9e6d8e60f7d9b835af06c75314399c8572a3b5cf38bceb56
|
File details
Details for the file podcast_helper-1.1.5-py3-none-any.whl.
File metadata
- Download URL: podcast_helper-1.1.5-py3-none-any.whl
- Upload date:
- Size: 44.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
db20a1c8122505f836e77ee0598cc927a5100ef867be6c929f0ab383778c6689
|
|
| MD5 |
2041a0a287fd40ae414bd35c3dc9b4bd
|
|
| BLAKE2b-256 |
5b5e023aeab035f5113ad3b9662c6f6f9fe62010ab8ac15dfe737c93ccacc3a0
|