SPORC: Structured Podcast Open Research Corpus
A Python package for working with the SPoRC dataset. It covers 228,099 podcasts and 1,124,058 episodes of transcripts, speaker turns, and acoustic features.
The dataset itself (schema, columns, layout, terms of use) is documented on the dataset card. This README covers the Python package.
📖 Full documentation: sporc.readthedocs.io, with guides and a complete API reference. 🎓 Tutorials: eight worked notebooks built around real research questions.
Installation
The dataset is gated, so before installing:
- Accept the terms at huggingface.co/datasets/blitt/SPoRC. Log in and click "Agree".
- Authenticate locally:
pip install huggingface_hub hf auth login
(On olderhuggingface_hubthis command washuggingface-cli login.)
Then:
pip install sporc
Or from source:
git clone https://github.com/davidjurgens/sporc.git
cd sporc
pip install -e .
Full-text search also needs DuckDB: pip install sporc[duckdb].
Quick Start
from sporc import SPORCDataset
# Downloads the metadata catalogs (~195 MB). Episode and turn data is fetched
# per podcast, as you touch it.
sporc = SPORCDataset()
# search_podcast matches by exact title first, then falls back to a substring
# match — so prefer the full, exact title (or a podcast_id) to avoid surprises.
podcast = sporc.search_podcast("The NPR Politics Podcast")
print(podcast.title, podcast.num_episodes, "episodes")
for episode in podcast.episodes:
print(f"{episode.title} — {episode.duration_minutes:.0f} min")
# Only ~65% of episodes were diarized; check before using turns.
if episode.has_turn_data:
for turn in episode.turns[:3]:
print(f" [{turn.inferred_speaker_role}] {turn.text[:60]}...")
Data Access
The corpus is ~57 GB and partitioned by podcast, so each podcast is fetched as a unit. A median podcast is ~75 KB, so the one-time ~195 MB of catalogs dominates any study of less than a few thousand podcasts:
| Study size | Data fetched |
|---|---|
| 10 podcasts | ~750 KB |
| 200 podcasts | ~15 MB |
| 1,000 podcasts | ~75 MB |
| whole corpus | ~31 GB |
# Default: catalogs now, per-podcast data on demand.
sporc = SPORCDataset()
# Fetch a known slice up front. Accepts podcast ids, podcast titles, episode
# ids, or a path to a .json / newline-delimited .txt file of them.
sporc = SPORCDataset(subset=["The NPR Politics Podcast", "Stuff You Should Know"])
sporc = SPORCDataset(subset="my_podcast_ids.txt")
# Pin a run to that slice: anything outside it raises DataNotLocalError
# rather than quietly downloading more.
sporc = SPORCDataset(subset="my_ids.txt", allow_downloads=False)
# Use a local copy of the layout; never downloads anything.
sporc = SPORCDataset(parquet_dir="/path/to/sporc_parquet")
# Download the whole corpus up front (~31 GB, ~685k files).
sporc = SPORCDataset(lazy=False)
prefetch() does the same job after construction, and reports what it resolved:
sporc.prefetch(["The NPR Politics Podcast"])
# {'podcasts': 1, 'files': 4, 'unresolved': []}
Selecting a subset efficiently
Selection is metadata-only. The calls below run off the already-downloaded catalogs and fetch nothing, so you can narrow to exactly the episodes you want and only then pull data:
hits = sporc.filter_episodes_by_metrics(min_word_count=5000, limit=200)
hits = sporc.search_by_speaker_name("Ira Glass", role="host")
sporc.prefetch({"episode_ids": [h["episode_id"] for h in hits]})
Two things worth knowing:
- Pass
max_episodestosearch_episodes. Matching is metadata-only, but building eachEpisodereads that podcast's partition.category="comedy"matches 62,622 episodes across 14,668 podcasts, somax_episodes=10reads 10 partitions instead. filter_episodes_by_metricsimplies turn data.episode_metricsis derived from turns, so it only covers diarized episodes and never fetches one without them. Filtering the catalog directly,num_main_speakers > 0marks the same set.
Time span
SPoRC is a two-month snapshot: every episode was published between 1 May and 30 June 2020 (median 28 May). It is not a longitudinal archive, so "trends over time" means trends across eight weeks. The window does straddle a sharp, dateable event, which makes before/after designs unusually clean.
Turn coverage
Only 731,101 of 1,124,058 episodes (65%) were diarized into speaker turns.
The rest have a transcript but are marked SPEAKER_DATA_UNAVAILABLE upstream.
Dataset version 1.1 roughly doubled this: 1.0 had 372,604 (33%), and 358,497
episodes that had been diarized but never merged were added.
An empty episode.turns is therefore usually a gap in the corpus rather than a
fact about the episode:
if episode.has_turn_data:
analyze(episode.turns)
Text search
search_turns, search_episodes_by_text and concordance use the DuckDB
full-text index when present, and otherwise scan the turn partitions on disk:
| Local data | Scan (pyarrow) | 26 GB DuckDB index |
|---|---|---|
| 250 episodes (5 MB) | 0.37s | n/a — needs the full 26 GB |
| 1,000 episodes (17 MB) | 1.1s | " |
| 5,000 episodes (83 MB) | 5.4s | " |
| all 78M turns | impractical | 1.7s open, ~5s/query (50s cold) |
The 26 GB index is only worthwhile for whole-corpus work; below ~5,000 episodes, scanning beats its cold start alone.
sporc = SPORCDataset(include_search_db=True) # whole corpus; pip install duckdb
results = sporc.search_turns("artificial intelligence") # fts (BM25)
results = sporc.search_turns("climate", mode="exact") # substring
results = sporc.concordance("like", context_words=5) # KWIC
Without the index, mode="fts" ranks by term frequency rather than BM25 and
covers only local data; the package warns and reports how many podcasts it
scanned.
Building teaching subsets
scripts/make_subset.py cuts a self-contained mini-SPoRC. It filters the
catalogs and episode partitions together so counts, searches and statistics
are all true of the subset:
# Ten disjoint 1k-episode subsets (~57 MB each)
for i in $(seq 1 10); do
python scripts/make_subset.py --data-dir /path/to/sporc_parquet \
--out subsets/subset_$i --episodes 1000 --seed $i \
--exclude-used subsets/used.txt
done
Subsets are diarized-only by default, so every episode has turns; pass
--include-undiarized to mirror the real corpus's ~65% coverage. Learners point
at the directory and nothing downloads:
sporc = SPORCDataset(parquet_dir="subsets/subset_1")
Tutorials
Eight standalone notebooks in examples/notebooks/, each
framed around a real research question. They run against a small pre-built
tutorial subset (not the full 57 GB corpus). See the
notebooks README for setup. Start with 01; it
establishes the corpus caveats the rest depend on.
| # | Notebook | Question |
|---|---|---|
| 01 | Corpus cartography | What is actually in SPoRC? |
| 02 | NER co-mention networks | Who gets talked about together? |
| 03 | Host/guest networks | Which shows share guests? |
| 04 | Repeat-guest language | Do repeat guests reuse material? |
| 05 | Stance over time | How did talk change around 25 May 2020? |
| 06 | Topic modeling (MALLET) | What is podcasting about? |
| 07 | Sociophonetics: caught/cot | Does this speaker merge caught and cot? |
| 08 | Conversational dynamics | Who talks longest, and where do people overlap? |
If you only read one, read 07. It starts from a question the corpus cannot answer and shows what to do about it.
Shorter, single-topic example scripts live in examples/
(basic_usage.py, category_examples.py, sliding_window_examples.py,
time_range_examples.py, advanced_analysis.py).
Core Classes
SPORCDataset
Search and retrieval over the corpus.
sporc.search_podcast("The NPR Politics Podcast") # -> Podcast
sporc.search_episodes(min_duration=1800, category="science", max_episodes=20)
sporc.search_episodes_by_subcategory("Technology", max_episodes=10)
sporc.get_dataset_statistics() # counts, distributions
for episode in sporc.iterate_episodes(max_episodes=100, sampling_mode="random"):
...
# Precomputed metrics (diarized episodes only)
sporc.get_episode_metrics(episode_id)
sporc.filter_episodes_by_metrics(min_turn_count=50, max_host_proportion=0.6)
sporc.get_turn_metrics(podcast_id, episode_id)
Podcast
podcast.title, podcast.num_episodes, podcast.primary_category
podcast.total_duration_hours, podcast.avg_episode_duration_minutes
podcast.host_names, podcast.categories
podcast.get_episodes_by_host("Jad Abumrad")
podcast.get_episodes_by_duration_range(600, 1800)
podcast.interview_episodes, podcast.solo_episodes, podcast.panel_episodes
Episode
episode.title, episode.duration_minutes, episode.transcript
episode.host_names, episode.guest_names, episode.num_main_speakers
episode.primary_category, episode.categories
episode.has_turn_data # False when the corpus has no turns for it
episode.turns # lazily loaded
episode.turn_count
episode.get_turns_by_time_range(0, 180)
episode.get_host_turns(), episode.get_guest_turns()
episode.get_turn_statistics()
Turn
turn.text, turn.speaker, turn.duration
turn.start_time, turn.end_time
turn.inferred_speaker_role # "host" | "guest" | "NO_INFERRED_ROLE"
turn.inferred_speaker_name
turn.word_count, turn.words_per_second
turn.is_host, turn.is_guest
turn.get_audio_features() # MFCCs, F0, F1
The acoustics are thin. get_audio_features() returns six numbers
(mfcc1..4_sma3_mean, one F0 mean, one F1 mean), each averaged over the whole
turn. There is no F2, no frame-level contour, and no word-level timing anywhere
in SPoRC. Vowel-level questions cannot be answered from these fields; see
Phonetics for the way around it.
Role labels are sparse. In a 174k-turn sample, 90.6% of turns are
NO_INFERRED_ROLE, with only 7.4% host and 1.9% guest. So get_host_turns(),
search_turns(speaker_role=...) and role distributions cover a small slice of an
episode, and an unlabelled turn does not mean nobody spoke. Treat role-based
counts as a lower bound rather than a partition of the conversation.
Sliding Windows
Process long episodes in chunks, by turn count or by time:
for window in episode.sliding_window(window_size=10, overlap=2):
print(f"{window.size} turns, {window.duration/60:.1f} min")
print(window.get_text())
print(window.get_role_distribution())
for window in episode.sliding_window_by_time(window_duration=300, overlap_duration=60):
...
Phonetics
SPoRC has no word timings, and its acoustics are six per-turn means with no F2,
so vowel-level work is impossible from the corpus alone. But every turn carries an
mp3_url plus its own start_time/end_time, which is enough to fetch that turn
and derive alignment properly.
sporc.phonetics does that. It is optional and its dependencies are heavy, so it
is opt-in and imported lazily:
pip install sporc[phonetics] # torch, torchaudio, transformers, parselmouth
# plus an ffmpeg binary on PATH
from sporc.phonetics import find_word_tokens, lobanov_normalize
# search -> fetch turn audio -> forced-align -> measure the vowel
tokens = find_word_tokens(sporc, "caught", limit=50)
df = lobanov_normalize(tokens) # per-speaker z-scores; raw formants
# mostly measure vocal-tract length
Lower-level pieces, if you want the steps:
from sporc.phonetics import (fetch_turn_audio, align_turn, measure_formants,
stressed_vowel_index)
audio, sr = fetch_turn_audio(turn) # range-fetches ONLY the turn
words = align_turn(turn, audio=audio, sample_rate=sr, level="word")
# level="phone" aligns one word's clip, not a whole turn: forced alignment makes
# the phones explain all the audio you hand it, so slice the word out first.
hit = next(w for w in words if w.word.lower() == "caught")
clip = audio[int(hit.start * sr):int(hit.end * sr)]
phones = align_turn(turn, audio=clip, sample_rate=sr, level="phone", word="caught")
# Ask which phone is the stressed vowel rather than counting positions: it is
# the second phone in "caught", but the third in "across".
v = phones[stressed_vowel_index([p.arpabet for p in phones])]
measure_formants(clip, sr, v.start, v.end) # F1/F2/F3
Worth knowing before you rely on it:
- Only the turn is downloaded. ffmpeg seeks into the remote mp3 with an HTTP range request: ~1–2 s and a few hundred KB per turn, not a 100 MB episode.
- Aligning costs more than fetching. Finding a word means aligning the whole
turn, at ~0.45x realtime on CPU, so cost scales with the turn, not the word.
Turns run long (median ~64 s for a common word, longest in the corpus 3,240 s),
so
find_word_tokensskips turns overmax_turn_duration=30by default and logs what it dropped. Alimit=50call takes minutes, not seconds. - The audio is external.
mp3_urlpoints at the publisher's CDN, not HuggingFace. Links rot, and some hosts need redirects resolved first, whichfetch_turn_audiodoes. - A speaker is identified only by name.
lobanov_normalizegroups by inferred name across episodes and shows, and dropsNO_INFERRED_SPEAKERand rawSPEAKER_00labels rather than pooling them. The placeholder is the corpus's most common speaker value and would otherwise average unrelated people into one voice. Two real people sharing a name still pool, because SPoRC has no canonical speaker id. estimate_word_audio()is not a substitute for this. It interpolates by character offset and is deprecated for phonetic use.
examples/notebooks/07_sociophonetics_caught_cot.ipynb works the whole thing
through on the caught/cot merger.
Error Handling
from sporc import SPORCDataset, SPORCError, NotFoundError, DataNotLocalError
try:
sporc = SPORCDataset(subset="ids.txt", allow_downloads=False)
podcast = sporc.search_podcast("Some Podcast")
except NotFoundError:
... # no such podcast in the catalog
except DataNotLocalError:
... # needed data is not local and downloads are disabled
except SPORCError:
... # base class for all package errors
Development
pip install -e ".[dev]"
pytest # full suite
pytest -m "not slow and not integration"
pytest --cov=sporc
black sporc/ tests/
isort sporc/ tests/
flake8 sporc/ tests/
mypy sporc/
Citation
If you use SPoRC in your research, please cite:
@inproceedings{litterer-etal-2025-mapping,
title = "Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus",
author = "Litterer, Benjamin Roger and Jurgens, David and Card, Dallas",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1222/",
doi = "10.18653/v1/2025.acl-long.1222",
pages = "25132--25154",
}
License
MIT; see LICENSE. The dataset itself is released for research and educational use only; see the dataset card for its terms.
Support
- Documentation: sporc.readthedocs.io
- Tutorials:
examples/notebooks/ - Questions, bugs, feature requests: open an issue
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sporc-1.1.4.tar.gz.
File metadata
- Download URL: sporc-1.1.4.tar.gz
- Upload date:
- Size: 133.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.11.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7de81fa1ec0acc2971f919a2288f1096d42e4b4326095a4f22e24d340c1674fa
|
|
| MD5 |
43432aa7ec758315e14eec3b0330d1ab
|
|
| BLAKE2b-256 |
7917cbcbc61912316ad0ddb258fd0e016250dfd6b73d738b8a57ba5c9223706a
|
File details
Details for the file sporc-1.1.4-py3-none-any.whl.
File metadata
- Download URL: sporc-1.1.4-py3-none-any.whl
- Upload date:
- Size: 84.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.11.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e19c87fc707a260db4b67ddf77b62c2d6dfd7de5b7b6f9258ae1e6d6c5ba8565
|
|
| MD5 |
defb033acae90eabd844d7dc51edbf79
|
|
| BLAKE2b-256 |
95ab80a857064f208c1554becffb1356a92253f64f67b0a93f343606544eff9b
|