Skip to main content

YouTube Scrapper

The YouTube scrapper is a data pipeline that allows you to extract data from videos posted on YouTube. The scrapper will extract information such as; video titles, descriptions, the number of views the video has and all the users or viewers' comments per each video. Information extracted is dumped into a csv file which can then be uploaded to an Amazon RDS database.

Getting Started

  1. Make sure pip and Chromedriver is installed.
  2. Download package by using pip install Youtube-scrapper in console
  3. Run 'python -m youtube_scrapper' on your terminal
  4. A prompt to enter the following information will pop up;
    1. channel - this takes the youtube channel you want to scrape data from
    2. scroll_range - this indicates how many video you want the scrapper to look and grab information from.
  5. Once the information is passed in, the scrapper will start grabbing all information requested.
  6. Once finished, the data is returned in a csv file which can be found at the directory level the package was run.

Note

  • When the scroll_range value is 0, 30 videos (The least amount of videos that will be scrapped) are returned from the channel specified. Each time the value goes up by 1, 30 more videos are returned.
  • Run time takes a while (approximately 30mins for 30 videos)

Dependencies

  • Python 3.7.9
  • selenium 3.141.0
  • pandas 1.3.2

Release files for Youtube-scrapper 0.0.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for Youtube-scrapper 0.0.7
File Size Uploaded
Youtube_scrapper-0.0.7.tar.gz 5.4 kB Details

Release files / Youtube_scrapper-0.0.7.tar.gz

Download URL Youtube_scrapper-0.0.7.tar.gz
Size 5.4 kB
Tags Source
SHA-256 checksum
How to use checksums
3282883b2c4ed885a261b83fe528d7b007b9619e4b3e627461947b01c9ab29a1
BLAKE2b-256 checksum
How to use checksums
3af6e149f71b226420233f49ea90ccd9c6fa07d43db922956e9ebcccb5d60351
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.2 importlib_metadata/4.8.1 pkginfo/1.7.1 requests/2.26.0 requests-toolbelt/0.9.1 tqdm/4.61.2 CPython/3.9.5

Release history Release notifications | RSS feed

This release

0.0.7 This release

1 release file

0.0.5

1 release file

0.0.4

1 release file

0.0.3

1 release file

0.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page