Skip to main content

YouTube Transcript API: Python SDK

License Website Python

The official Python SDK (getyoutubetranscript) for the GetYouTubeTranscript YouTube Transcript API. Get YouTube video transcripts, captions and subtitles (optionally with per-line timestamps) in Python without a Google API key, yt-dlp, or a headless browser. Get YouTube transcripts, search videos and channels, resolve channel handles, browse a channel's full upload history, search inside a channel, pull playlist contents, and check your credit balance, all with one typed client.

Export transcripts as plain text, timed text, JSON, SRT or WebVTT, from Python or the getyoutubetranscript command line. Getting RequestBlocked or IpBlocked from youtube-transcript-api on a cloud server? See below.

PyPI

Install

pip install getyoutubetranscript

Requires Python 3.9+.

Quickstart

from getyoutubetranscript import Client

client = Client(api_key="sk_live_...")

transcript = client.get_transcript("https://www.youtube.com/watch?v=jNQXAC9IVRw")
print(transcript["title"], transcript["word_count"])
print(transcript["transcript"])

Getting an API key

Every request needs an API key. There are two ways to get one:

  1. Dashboard - sign up at getyoutubetranscript.com. Free tier: 100 credits, no card required.
  2. Self-serve, in code - use the signup/verify_signup helpers below. No key required for either call.
from getyoutubetranscript import signup, verify_signup

signup("you@example.com")          # sends a 6-digit code, valid 10 minutes
# ... read the code from your inbox ...
api_key = verify_signup("you@example.com", "123456")  # -> "sk_live_..."

The raw key is returned once by verify_signup and can't be retrieved again - store it yourself (env var, secret manager, etc).

Usage

Every method costs 1 credit unless noted "free" below. Failed and rate-limited requests are never charged. All methods raise GetYouTubeTranscriptError on failure - see Error handling.

Transcripts

client.get_transcript("jNQXAC9IVRw", language="en")

Pass timestamps=True to also get one entry per caption line in segments (same 1 credit). Without it, the response has no segments key.

result = client.get_transcript("5e37ZT3SQbk", timestamps=True)
print(result["segments"][0])
# {"start": 3.96, "duration": 4.56, "text": "So, Reed, education, which a lot of"}

Each segment is {"start", "duration", "text"} with start and duration in seconds. The Segment and TranscriptData typed dicts are importable from getyoutubetranscript.

Every transcript also says where it came from:

Field Meaning
language_code The caption track actually returned
requested_language What you asked for. If it differs from language_code, YouTube didn't have that language
caption_type "manual" (uploaded by the creator), "auto" (YouTube speech recognition), or None if unknown
cached True when served from the stored copy rather than fetched from YouTube just now
fetched_at ISO 8601 time it was fetched from YouTube

Batch: many videos at once

Queue up to 100 videos in one call; transcripts are fetched in the background. Submitting is free, each video that returns a transcript costs 1 credit, and failed videos are never charged (10 videos where 2 have no captions = 8 credits).

batch = client.create_batch(["jNQXAC9IVRw", "https://youtu.be/dQw4w9WgXcQ"], language="en")
result = client.wait_for_batch(batch["batch_id"])  # polls, then collects every page

for item in result["items"]:
    if item["status"] == "succeeded":
        print(item["video_id"], item["caption_type"], item["transcript"][:80])
    else:
        print(item["video_id"], "failed:", item["error_code"])  # e.g. TRANSCRIPT_DISABLED
print("credits used:", result["credits_charged"])

get_batch(batch_id, offset=0, limit=20) returns the status and one page if you'd rather poll yourself. Pass idempotency_key="..." to create_batch so a retried call returns the same batch instead of queuing a second one. Batches are kept for 7 days, and an account can have 5 unfinished batches at a time.

To be notified instead of polling, pass webhook_url (public https). The response includes a webhook_secret (shown once); each delivery is signed, so check it before trusting the body:

from getyoutubetranscript import verify_webhook_signature

# Flask example: use the raw body, not re-serialized JSON
@app.post("/hooks/transcripts")
def transcripts_done():
    if not verify_webhook_signature(request.get_data(), request.headers.get("X-GYT-Signature"), WEBHOOK_SECRET):
        abort(401)
    event = request.get_json()  # {"event": "batch.completed", "batch_id": ..., "results_url": ...}
    ...

Formats: text, timed text, JSON, SRT, WebVTT

Turn a transcript into a file format with the formatters. Timed text, SRT and WebVTT need per-line timing, so fetch with timestamps=True (same 1 credit).

from getyoutubetranscript import Client, to_srt, to_vtt, to_timed_text, to_json, to_text

result = client.get_transcript("5e37ZT3SQbk", timestamps=True)

open("video.srt", "w", encoding="utf-8").write(to_srt(result))   # SubRip subtitles
open("video.vtt", "w", encoding="utf-8").write(to_vtt(result))   # WebVTT subtitles
print(to_timed_text(result))  # "[0:03] So, Reed, education, ..." one line per caption
print(to_json(result))        # metadata, transcript and segments
print(to_text(result))        # one block of plain text

format_transcript(result, "srt") does the same with the format as a string ("text", "timed", "json", "srt", "vtt").

first_page = client.search("lofi beats", type="video", limit=10)

# Pagination: pass continuation_token back as page_token
if first_page.get("continuation_token"):
    page2 = client.search(page_token=first_page["continuation_token"])

Channels

client.resolve_channel("@mkbhd")            # free - handle/URL -> channel ID
client.get_channel_latest("@mkbhd")         # free - metadata + latest uploads
client.search_channel("@mkbhd", "iphone")   # search within a channel
client.list_channel_videos("@mkbhd")        # full paginated upload history

# Pagination (search_channel and list_channel_videos both work the same way)
page = client.list_channel_videos("@mkbhd")
while page["has_more"]:
    page = client.list_channel_videos(continuation=page["continuation_token"])

Playlists

page = client.get_playlist("PLillGF-RfqbYE6Ik_EuXA2iZFcE082B3s")
while page["has_more"]:
    page = client.get_playlist(continuation=page["continuation_token"])

Account

client.get_credits()  # free - plan_credits_left, topup_credits_left, plan, rate_limit_per_minute

Command line

Installing the package also installs a getyoutubetranscript command.

export GETYOUTUBETRANSCRIPT_API_KEY=sk_live_...

getyoutubetranscript https://youtu.be/5e37ZT3SQbk                 # plain text
getyoutubetranscript 5e37ZT3SQbk --format srt > video.srt          # SRT subtitles
getyoutubetranscript 5e37ZT3SQbk --format timed --language en      # [m:ss] lines
getyoutubetranscript VIDEO_1 VIDEO_2 --format json                 # one JSON list
getyoutubetranscript VIDEO_1 VIDEO_2 --format vtt --output-dir subs # subs/<video_id>.vtt

Formats: text (default), timed, json, srt, vtt. Videos can be URLs or IDs. If one video fails, the rest still run and the command exits with status 1. python -m getyoutubetranscript works too.

Error handling

Every non-2xx or {"success": false} response raises GetYouTubeTranscriptError with the API's parsed error shape:

from getyoutubetranscript import Client, GetYouTubeTranscriptError

client = Client(api_key="sk_live_...")

try:
    client.get_transcript("no-captions-here")
except GetYouTubeTranscriptError as e:
    print(e.code)          # e.g. "NOT_FOUND"
    print(e.message)       # human-readable message from the API
    print(e.status_code)   # 400 / 401 / 402 / 404 / 429 / 503, or 0 for a local network failure
    print(e.response_body) # full parsed error body, e.g. {"creditsLeft": 0} on PAYMENT_REQUIRED

Getting RequestBlocked or IpBlocked?

If you use the open source youtube-transcript-api library, you have probably seen RequestBlocked or IpBlocked once your code runs on a server. YouTube blocks most IP addresses that belong to cloud providers (AWS, Google Cloud, Azure and others), and can also block a home IP that makes many requests. That library's own docs recommend rotating residential proxies as the workaround.

This SDK calls the GetYouTubeTranscript API instead of YouTube, so YouTube never sees your server's IP. There are no proxies to buy, rotate or debug, and the same code works on your laptop, a VPS, a serverless function or a CI job:

from getyoutubetranscript import Client

client = Client(api_key="sk_live_...")
result = client.get_transcript("https://www.youtube.com/watch?v=jNQXAC9IVRw", timestamps=True)

The trade-off: it is a paid API with a free tier (each request uses credits), while youtube-transcript-api is free to run if you handle the blocking yourself.

Coming from youtube-transcript-api

The segment shape is the same (text, start, duration, in seconds), so most code ports directly.

youtube-transcript-api getyoutubetranscript
YouTubeTranscriptApi().fetch(video_id) client.get_transcript(video, timestamps=True)
fetched.to_raw_data() result["segments"] (already a list of dicts)
fetch(video_id, languages=["de"]) get_transcript(video, language="de")
Video ID only Video ID or any YouTube URL (watch, youtu.be, Shorts, live)
SRTFormatter(), WebVTTFormatter(), TextFormatter(), JSONFormatter() to_srt, to_vtt, to_text, to_json
CLI: youtube_transcript_api VIDEO_ID --format json CLI: getyoutubetranscript VIDEO --format json
Proxies for cloud servers Not needed
# before
from youtube_transcript_api import YouTubeTranscriptApi
segments = YouTubeTranscriptApi().fetch("jNQXAC9IVRw").to_raw_data()

# after
from getyoutubetranscript import Client
segments = Client(api_key="sk_live_...").get_transcript("jNQXAC9IVRw", timestamps=True)["segments"]

Not covered here: a list of preferred fallback languages, listing every available caption track, YouTube's machine translation of captions, and preserve_formatting. Request one language at a time with language=.

The response also includes the video title, channel name, channel URL, thumbnail and word count, which youtube-transcript-api does not return.

Development

pip install -e ".[dev]"

# Unit tests - mocked HTTP, no network or API key needed, always safe to run
pytest tests -v --ignore=tests/live

# Live integration tests - hits the real API, spends credits, needs a key
GYT_API_KEY=sk_live_... pytest tests/live -v

Other ways to use the GetYouTubeTranscript API:

License

MIT - see LICENSE.

Metadata

Release files for getyoutubetranscript 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for getyoutubetranscript 0.4.0
File Size Uploaded
getyoutubetranscript-0.4.0.tar.gz 24.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for getyoutubetranscript 0.4.0
File Interpreter ABI Platform
getyoutubetranscript-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 46.1 kB

Release files / getyoutubetranscript-0.4.0.tar.gz

Download URL getyoutubetranscript-0.4.0.tar.gz
Size 24.6 kB
Tags Source
SHA-256 checksum
How to use checksums
643768ab3fca6781ee2071d2ba9d0eb1fa9bb1ee861eb4a0e324632e6cc024dc
BLAKE2b-256 checksum
How to use checksums
5f12d63b013a5f3c6d7957929064172384d0d31e32c25cf09e5deb4ebb669e8c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.0

Release files / getyoutubetranscript-0.4.0-py3-none-any.whl

Download URL getyoutubetranscript-0.4.0-py3-none-any.whl
Size 21.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
04bf7cd0f3da053fff8e6693e3dab9d71cba621ac4b1a42b1eb60cf125047b4d
BLAKE2b-256 checksum
How to use checksums
973795833d16960fb99698450000c8e04eeb41bd23296431c854b63ab08547d4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.0

Release history Release notifications | RSS feed

0.5.0

2 release files

This release

0.4.0 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page