Skip to main content

getyoutubetranscript-haystack

A Haystack component that fetches YouTube video transcripts as Documents, powered by the GetYouTubeTranscript API. Use it in RAG pipelines, or wrap it as a tool for Haystack agents.

The API fetches transcripts on its own servers, so it works from cloud servers without proxies, and without RequestBlocked / IpBlocked errors.

Install

pip install getyoutubetranscript-haystack

Get an API key at getyoutubetranscript.com/dashboard (free tier included):

export GETYOUTUBETRANSCRIPT_API_KEY=sk_live_...

Usage

from haystack_integrations.components.fetchers.getyoutubetranscript import GetYouTubeTranscriptFetcher

fetcher = GetYouTubeTranscriptFetcher()
documents = fetcher.run(videos=["https://youtu.be/jNQXAC9IVRw", "5e37ZT3SQbk"])["documents"]

print(documents[0].meta)
# {'url': 'https://www.youtube.com/watch?v=jNQXAC9IVRw', 'video_id': 'jNQXAC9IVRw', 'title': 'Me at the zoo',
#  'author_name': 'jawed', 'language_code': 'en', 'word_count': 39}

One Document per video. Parameters:

Parameter Default Description
api_key Secret.from_env_var("GETYOUTUBETRANSCRIPT_API_KEY") API key
language None Caption language code, e.g. "en"
timestamps False content becomes [m:ss] lines, and meta["segments"] holds {start, duration, text} per caption line
raise_on_failure True If False, videos that fail (no captions, invalid ID) are logged and skipped

In a RAG pipeline

from haystack import Pipeline
from haystack.components.preprocessors import DocumentSplitter
from haystack.components.writers import DocumentWriter
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack_integrations.components.fetchers.getyoutubetranscript import GetYouTubeTranscriptFetcher

store = InMemoryDocumentStore()
indexing = Pipeline()
indexing.add_component("fetcher", GetYouTubeTranscriptFetcher())
indexing.add_component("splitter", DocumentSplitter(split_by="word", split_length=200))
indexing.add_component("writer", DocumentWriter(document_store=store))
indexing.connect("fetcher.documents", "splitter.documents")
indexing.connect("splitter.documents", "writer.documents")

indexing.run({"fetcher": {"videos": ["https://youtu.be/5e37ZT3SQbk"]}})

As an agent tool

from haystack.tools import ComponentTool
from haystack_integrations.components.fetchers.getyoutubetranscript import GetYouTubeTranscriptFetcher

youtube_tool = ComponentTool(
    component=GetYouTubeTranscriptFetcher(timestamps=True),
    name="youtube_transcript",
    description="Get the transcripts of YouTube videos from their URLs or IDs.",
)

The component serializes with to_dict() / from_dict(), keeping the key as an environment variable reference, so pipelines can be saved as YAML.

Pricing

Each transcript uses one credit from your GetYouTubeTranscript account. Failed requests are not charged.

Metadata

Release files for getyoutubetranscript-haystack 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for getyoutubetranscript-haystack 0.1.0
File Size Uploaded
getyoutubetranscript_haystack-0.1.0.tar.gz 5.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for getyoutubetranscript-haystack 0.1.0
File Interpreter ABI Platform
getyoutubetranscript_haystack-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 11.2 kB

Release files / getyoutubetranscript_haystack-0.1.0.tar.gz

Download URL getyoutubetranscript_haystack-0.1.0.tar.gz
Size 5.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f308827deaeabd30a0a3788d2f4751e52d397c1b8dadece7793a0445013bd665
BLAKE2b-256 checksum
How to use checksums
80d9f04c42290c17b6e214bd9c98b626195596c9b92768c94d8bda6136a4cb8a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.0

Release files / getyoutubetranscript_haystack-0.1.0-py3-none-any.whl

Download URL getyoutubetranscript_haystack-0.1.0-py3-none-any.whl
Size 5.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e1398e8c59dde6cd7481c3d38848741529d325c67681490d8e1ca0d8af46d474
BLAKE2b-256 checksum
How to use checksums
f28ca82a41edfc117b366a33e7043bd66da8c147dc28d70929cdc785abde5603
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.0

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page