Skip to main content

YouTube Channel Scraper

Scrapes videos from a YouTube channel using Python.

It takes a YouTube channel URL as an input and produces a list of videos as an output.

Video attributes:

  • url
  • title
  • description
  • author
  • published_at
  • thumbnail

Requirements

Install chromedriver at /chromedriver. For more details refer to scripts/install_chromedriver.sh.

Alternatively you can use the library inside a Docker container. This way you don't need to manually install the chromedriver. Have a look at Usage with Docker.

Install

pip install youtube-channel-scraper

Usage

from youtube_channel_scraper.scraper import YoutubeScraper

scraper = YoutubeScraper(channel_url)
crawled_videos = scraper.scrape()

Optional parameters

  • stop_id: stop scraping when this video ID is encountered. Older videos are always scraped first.
  • max_videos: maximum amount of videos to scrape
  • proxy_ip: proxy IP with port

CLI usage

youtube-channel-scraper "https://www.youtube.com/channel/CHANNEL_ID"

Available arguments

positional arguments:
  channel_url           Youtube channel URL

optional arguments:
  -h, --help            show this help message and exit
  --filename FILENAME   Output file
  --stop_video_id STOP_VIDEO_ID
                        Youtube video ID that indicates to stop crawling
  --max_videos MAX_VIDEOS
                        Max videos to crawl
  --proxy PROXY         Proxy IP
  --screenshot_filename SCREENSHOT_FILENAME
                        Path to exceptions screenshot

Usage with Docker

Why?

When running youtube-channel-scraper in a Docker container there is no need to manually install chrome driver in your environment.

How?

git clone https://github.com/hajkr/youtube-channel-scraper.git
cd youtube_channel_scraper

# Run the container and build it when running for the first time
make run

# Enter the container
make to_container

python youtube_channel_scraper "https://www.youtube.com/channel/CHANNEL_ID"

Metadata

Release files for youtube-channel-scraper 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for youtube-channel-scraper 0.1.3
File Size Uploaded
youtube-channel-scraper-0.1.3.tar.gz 4.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for youtube-channel-scraper 0.1.3
File Interpreter ABI Platform
youtube_channel_scraper-0.1.3-py2.py3-none-any.whl Python 3, Python 2 none any Details

Total release size: 10.2 kB

Release files / youtube-channel-scraper-0.1.3.tar.gz

Download URL youtube-channel-scraper-0.1.3.tar.gz
Size 4.7 kB
Tags Source
SHA-256 checksum
How to use checksums
904b26f27b8739e19376e1a7d4b95793d1e36cb0768b5f3c410834305ec9a3b9
BLAKE2b-256 checksum
How to use checksums
d864f43ea615dd000a1117d3da37f2ccad2c7931005cc6fb2a3c6dd229eb35d9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.1.6 CPython/3.9.5 Linux/5.10.25-linuxkit

Release files / youtube_channel_scraper-0.1.3-py2.py3-none-any.whl

Download URL youtube_channel_scraper-0.1.3-py2.py3-none-any.whl
Size 5.5 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
e27f3f99ca6a862dff3cdcf3df4e18cdbba53da318315dfba36d98c9619650f1
BLAKE2b-256 checksum
How to use checksums
dfc46e3fe1cd91c946972bd098dda6479e0950efb2211ee8a9aed1b491680192
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.1.6 CPython/3.9.5 Linux/5.10.25-linuxkit

Release history Release notifications | RSS feed

This release

0.1.3 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.41

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page