Skip to main content

some web scraping tool for Weibo

Project description

This library is something I wrote to make my code look neat. These are some scraping functions for Weibo

1. scroll(scroll_num): to scroll down some pages type(scroll_num) = int

2. scroll_little_bit(scroll_num) :to scroll down a little bit, no more than one page, at a time, 250 pixel to be precise type(scroll_num) = int

  1. login_mobile(): to log in to weibo mobile web page

  2. login_web(): to log in to weibo web page

5. hot_people(hashtag):find the users who has made the most contribute to a certain topic area,on weibo web. Takes in a hashtag and return a data frame type(hashtag) = str

6. search_mobile(user_name): to search a user on weibo mobile page type(user_name) = str

7.repo_passive_name(scroll_num):to get some user’s name who likes to repost certain key user’s post,on weibo mobile web type(scroll_num) = int

8.repo_name_initiative(scroll_num):to got all the user’s names that the key user reposted from,on weibo mobile web type(scroll_num) = int

9.make_list_num_df(list_long):to make a dataframe from a list type(list_long) = list

10.clean_repo_name(repo_name_list):to get all the “None” out of the list type(repo_name_list) = list

11.get_content_and_time(scroll_num):to get some of the post of a key user and their post time on weibo mobile web type(scroll_num) = int

12.find_repo_in_content(content_list):to some of possible connections in post content type(content_list) = list

13.fan_num_mobile(user_name):to get the follower number of a certain user ,on weibo mobile web type(user_name) = str

14.fan_num_web(user_name):to get the follower number of a certain user ,on weibo web type(user_name) = str

15. get_repo_like_ratio(user_name,scroll_num):to get the repost/like ratio of a certain user ,on weibo mobile web type(user_name) = str type(scroll_num) = int

16.search_topic_general(hashtag, page_num):to search a hashtag and scrape all the content in the search result, store the scraped content in a dataframe type(hashtag) = str type(page_num) = int

17.search_keyword_general(hashtag, page_num):to search a keyword and scrape all the content in the search result, store the scraped content in a dataframe type(hashtag) = str type(page_num) = int

  1. Library created in 2022/6.27 ver 0.0.1

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

yiwen scraper-0.0.1.tar.gz (5.6 kB view details)

Uploaded Source

File details

Details for the file yiwen scraper-0.0.1.tar.gz.

File metadata

  • Download URL: yiwen scraper-0.0.1.tar.gz
  • Upload date:
  • Size: 5.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.1 CPython/3.8.12

File hashes

Hashes for yiwen scraper-0.0.1.tar.gz
Algorithm Hash digest
SHA256 d6c8d51d38120f512c9fb200bda49946e3133b5169460ed6f8e76351996e1e38
MD5 64f0162c092ce92ec1bffea5480b1d27
BLAKE2b-256 1452407ab6830a7f5fe8a2c5ce54d4f6a7379bd8e2312cac680cf86bfc3c3a56

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page