llama-index-readers-getyoutubetranscript
A LlamaIndex reader that loads YouTube video transcripts as Documents, powered by the GetYouTubeTranscript API.
It has the same load_data(ytlinks=[...]) call as YoutubeTranscriptReader, but transcripts are fetched by the API on its own servers. That means it keeps working on cloud servers, where YouTube blocks direct requests (RequestBlocked / IpBlocked), with no proxies to manage.
Install
pip install llama-index-readers-getyoutubetranscript
Get an API key at getyoutubetranscript.com/dashboard (free tier included):
export GETYOUTUBETRANSCRIPT_API_KEY=sk_live_...
Usage
from llama_index.readers.getyoutubetranscript import GetYouTubeTranscriptReader
reader = GetYouTubeTranscriptReader()
documents = reader.load_data(ytlinks=["https://youtu.be/jNQXAC9IVRw", "5e37ZT3SQbk"])
print(documents[0].metadata)
# {'url': 'https://www.youtube.com/watch?v=jNQXAC9IVRw', 'video_id': 'jNQXAC9IVRw', 'title': 'Me at the zoo',
# 'author_name': 'jawed', 'language_code': 'en', 'word_count': 39}
Switching from YoutubeTranscriptReader: change the import and class name; load_data(ytlinks=...) stays the same. Links can be any YouTube URL (watch, youtu.be, Shorts, live) or a video ID.
| Option | Default | Description |
|---|---|---|
api_key |
GETYOUTUBETRANSCRIPT_API_KEY env var |
API key (never serialized) |
language |
API default | Caption language code, e.g. "en" |
timestamps |
False |
text becomes [m:ss] lines; metadata["segments"] holds {start, duration, text} per line, hidden from embeddings and LLM prompts |
language and timestamps can also be passed per call: reader.load_data(ytlinks=[...], timestamps=True).
Build an index
from llama_index.core import VectorStoreIndex
from llama_index.readers.getyoutubetranscript import GetYouTubeTranscriptReader
documents = GetYouTubeTranscriptReader().load_data(ytlinks=["https://youtu.be/5e37ZT3SQbk"])
index = VectorStoreIndex.from_documents(documents)
print(index.as_query_engine().query("What does the speaker say about education?"))
Pricing
Each transcript uses one credit from your GetYouTubeTranscript account. Failed requests are not charged.
Links
- GetYouTubeTranscript API docs
- Python SDK (this package is built on it)
- License: MIT
Metadata
Release files for llama-index-readers-getyoutubetranscript 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llama_index_readers_getyoutubetranscript-0.1.0.tar.gz | 4.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llama_index_readers_getyoutubetranscript-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 10.2 kB
Release files / llama_index_readers_getyoutubetranscript-0.1.0.tar.gz
| Download URL | llama_index_readers_getyoutubetranscript-0.1.0.tar.gz |
|---|---|
| Size | 4.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
18d10078d9246ebabf8a557d706438bef851ecbcc432871d33e7ec2cbd14f88b
|
|
BLAKE2b-256 checksum How to use checksums |
d6e8d03a801b8bc93a971f1f71b97a8cf35bc504bdd81489f027e96f750e8f45
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.0
|
Release files / llama_index_readers_getyoutubetranscript-0.1.0-py3-none-any.whl
| Download URL | llama_index_readers_getyoutubetranscript-0.1.0-py3-none-any.whl |
|---|---|
| Size | 5.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
85a274cd83320bff0f9e246a816fee0f447e613b8f150bd34055a9e0525a185a
|
|
BLAKE2b-256 checksum How to use checksums |
dc6773a6974b3de0b20204fc7e150346a99e48d31ca317928d4d60d4845504bd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.0
|