Skip to main content

pg-podcast-toolkit

Tools for parsing and managing Podcasting 2.0 RSS feeds with automatic namespace capture and database-ready output.

Features

  • Podcasting 2.0 Support - Automatically captures all podcast:* namespace tags without parser updates
  • Database-Ready Output - Built-in to_db_record() methods for PostgreSQL schema alignment
  • Backward Compatible - Existing code continues to work, new features are opt-in
  • Deterministic IDs - MD5-based ID generation (32-char lowercase hex strings) for podcasts and episodes
  • GUID Fallback - Handles episodes with missing GUIDs gracefully
  • Comprehensive Parsing - Supports RSS 2.0, iTunes extensions, and custom namespaces

Installation

pip install pg-podcast-toolkit

Quick Start

Basic Usage

from pg_podcast_toolkit import Podcast
import requests

# Fetch and parse a podcast feed
response = requests.get('https://example.com/feed.xml')
podcast = Podcast(response.content, feed_url='https://example.com/feed.xml')

# Access podcast metadata
print(podcast.title)
print(podcast.description)
print(podcast.itunes_image)

# Access episodes
for item in podcast.items:
    print(f"{item.title} - {item.itunes_duration}s")

Database Integration (New in v0.2.0)

# Get database-ready podcast record
podcast_record = podcast.to_db_record(
    etag='some-etag',              # Optional HTTP ETag
    last_modified='Wed, 06 Nov',   # Optional Last-Modified header
    last_fetched_at=1234567890     # Optional fetch timestamp
)

# Insert into PostgreSQL
# podcast_record matches schema: id, podcast_guid, title, feed_url,
# image_url, language, itunes_id, etag, last_modified, last_fetched_at,
# created_at, updated_at, extras (JSONB)

# Get database-ready episode records
for item in podcast.items:
    episode_record = item.to_db_record(podcast_id=podcast_record['id'])
    # episode_record matches schema: id, podcast_id, guid, title,
    # description, image_url, publish_date, duration_seconds,
    # episode_number, season_number, episode_type, explicit,
    # enclosure_url, enclosure_type, enclosure_size,
    # created_at, updated_at, extras (JSONB)

Accessing Podcasting 2.0 Namespaces (New in v0.2.0)

# All unknown namespace tags are automatically captured
print(podcast.namespaces)
# {
#   'podcast': {
#     'guid': {'value': '...'},
#     'locked': {'value': 'yes', 'attributes': {'owner': 'email@example.com'}},
#     'funding': {'value': 'Support!', 'attributes': {'url': 'https://...'}},
#     'person': [
#       {'value': 'Host Name', 'attributes': {'role': 'host', 'img': '...'}},
#       ...
#     ]
#   }
# }

# Episode-level namespaces
for item in podcast.items:
    print(item.namespaces)
    # {
    #   'podcast': {
    #     'chapters': {'attributes': {'url': '...', 'type': 'application/json'}},
    #     'transcript': {'attributes': {'url': '...', 'type': 'text/srt'}},
    #     'person': [...],
    #     ...
    #   }
    # }

What's New in v0.2.0

  • Automatic Namespace Capture - No parser updates needed for new Podcasting 2.0 tags
  • Database-Ready Methods - Podcast.to_db_record() and Item.to_db_record()
  • Schema Alignment - Output matches PostgreSQL schema with MD5 hex string primary keys
  • GUID Fallback - Episodes without GUIDs use enclosure_url for ID generation
  • 100% Backward Compatible - All existing attributes and methods unchanged

Supported Specifications

  • RSS 2.0
  • iTunes Podcast Extensions
  • Podcasting 2.0 Namespace (automatic capture)
  • Custom namespace extensions (automatic capture)

Development Status

This library is actively maintained and production-ready. The v0.2.0 release introduces database integration features while maintaining full backward compatibility.

License

MIT License

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pg_podcast_toolkit-0.4.0.tar.gz (29.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pg_podcast_toolkit-0.4.0-py3-none-any.whl (20.2 kB view details)

Uploaded Python 3

File details

Details for the file pg_podcast_toolkit-0.4.0.tar.gz.

File metadata

  • Download URL: pg_podcast_toolkit-0.4.0.tar.gz
  • Upload date:
  • Size: 29.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pg_podcast_toolkit-0.4.0.tar.gz
Algorithm Hash digest
SHA256 c85bedd97950ab88f6f0297b9b656f7e8e5bbb496f0f7fc2a7757f5afdd6c2a9
MD5 fda01f0eb43df36e2c5c08e06214d14c
BLAKE2b-256 5a1421dbd6903b17f6a74284c731a12bce2c6f72b44716930f4ac90fa4d518ce

See more details on using hashes here.

File details

Details for the file pg_podcast_toolkit-0.4.0-py3-none-any.whl.

File metadata

File hashes

Hashes for pg_podcast_toolkit-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f15538d281df5ad754016c6ffb049588d61f5f33056bfb1a97b65ec8d25f3a00
MD5 1085ad4a23ccff652a2267fd584e4e27
BLAKE2b-256 a485462d3e9c2aa3d6f3d4c9aa897f3835add0c2001ffae04e386c0e45c457da

See more details on using hashes here.

Release history Release notifications | RSS feed

0.5.2

2 files

0.5.1

2 files

0.5.0

2 files

0.4.1

2 files

This release

0.4.0 This release

2 files

0.3.3

2 files

0.3.1

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page