Clark University, Package for YouTube crawler and cleaning data
Project description
clarku-youtube-crawler
Clark University YouTube crawler and JSON decoder for YouTube json. Please read documentation in DOCS
Installing
To install,
pip install clarku-youtube-crawler
The crawler needs multiple other packages to function.
If missing requirements (I already include all dependencies so it shouldn't happen), download requirements.txt
.
Navigate to the folder where it contains requirements.txt and run
pip install -r requirements.txt
Example usage
To initialize,
# your_script.py
import clarku_youtube_crawler as cu
test = cu.RawCrawler()
test.__build__()
test.crawl("searchkey",start_date=14, start_month=12, start_year=2020, day_count=2)
test.crawl_videos_in_list(comment_page_count=1)
test.merge_all()
channel = cu.ChannelCrawler()
channel.__build__()
channel.setup_channel(subscriber_cutoff=1000, keyword="")
channel.crawl()
channel.crawl_videos_in_list(comment_page_count=1)
channel.merge_all()
jsonn = cu.JSONDecoder()
jsonn.load_json("YouTube_RAW_20201221/FINAL_raw_merged.json")
Changelog
Version 0.0.1->0.0.3
This is beta without testing since python packaging is a pain. Please don't install these versions.
Version 0.0.5
Finally figured out testing. It works okay. More documentation to come.
Version 0.0.6
Stable release only for RawCrawler
feature
Version 1.0.0
Version 1.0.1
I think this might be our first full stable release.
Version 1.0.1.dev
Pre-release
Added different file types for ChannelCrawler. Added documentation
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Hashes for clarku_youtube_crawler-1.1.7.tar.gz
Algorithm | Hash digest | |
---|---|---|
SHA256 | 1aa23db20b5bbdd401141b0c5e5ae0554d34270bec4649e06d8515b4ab4e38b2 |
|
MD5 | 6bd83e62ba49760da84d1975e928bd28 |
|
BLAKE2b-256 | f7c1f9f5f2ca7d4797cd11c43d3b513c73cfc29044f20fd81b68a73e53cd7202 |
Hashes for clarku_youtube_crawler-1.1.7-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | dd9650cdcc60c4785a77a15783142d15a44181efcbb10f4326c0d682f6ac2640 |
|
MD5 | 58dd10058e2aa2cbaec4c8be5e1e87d7 |
|
BLAKE2b-256 | 22ba31659b71390e73c769ddd5d1730c44da9f3bba63f09d432f330d52ab1db2 |