Skip to main content

Reflow.

Mission.

Scrape data from reddit over a period of time of your choice, filter with AI assistants, and connect it to your ML pipelines!

Execution is as simple as this:

  • Make a config file with your required details of input.
  • Run the API in a single line with the config passed as input.

Installation.

pip install reflow

Docs.

1) Text API.

Argument Input Description
sort_by str Sort the results by available options like 'best', 'new' ,'top', 'controversial' , etc as available from Reddit.
subreddit_text_limit int Number of rows to be scraped per subreddit
total_limit int Total number of rows to be scraped
start_time DateTime Start date and time in dd.mm.yy hh.mm.ss format
stop_time DateTime Stop date and time in dd.mm.yy hh.mm.ss format
subreddit_search_term str Input search term to create filtered outputs
subreddit_object_type str Available options for scraping are submission and comment.
resume_task_timestamp str, Optional If task gets interrupted, the timestamp information available from the created folder names can be used to resume.

2) Image API

Argument Input Description
sort_by str Sort the results by available options like 'best', 'new' ,'top', 'controversial' , etc as available from Reddit.
subreddit_image_limit int Number of images to be scraped per subreddit
total_limit int Total number of images to be scraped
start_time DateTime Start date and time in dd.mm.yy hh.mm.ss format
stop_time DateTime Stop date and time in dd.mm.yy hh.mm.ss format
subreddit_search_term str Input search term to create filtered outputs
subreddit_object_type str Available options for scraping are submission and comment
client_id str Since Image API requires praw, the config requires a praw client ID.
client_secret str Praw client secret.

Examples

Text Scraping and filtering

config = {
        "sort_by": "best",
         "subreddit_text_limit": 50,
        "total_limit": 200,
        "start_time": "27.03.2021 11:38:42",
        "end_time": "27.03.2022 11:38:42",
        "subreddit_search_term": "healthcare",
        "subreddit_object_type": "comment",
         "resume_task_timestamp":1648613439
    }
from reflow import TextApi
TextApi(config)

Image Scraping and filtering

config = {
        "sort_by": "best",
        "subreddit_image_limit": 3,
        "total_limit": 10,
         "start_time": "13.11.2021 09:38:42",
         "end_time": "15.11.2021 11:38:42",
         "subreddit_search_term": "cats",
         "subreddit_object_type": "comment",
         "client_id": "$CLIENT_ID", # get client id for praw
         "client_secret": $CLIENT_SECRET, #get client secret for praw
         }

from reflow import ImageApi
ImageApi(config)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

redditflow-0.0.1.tar.gz (9.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

redditflow-0.0.1-py3.8.egg (28.8 kB view details)

Uploaded Egg

File details

Details for the file redditflow-0.0.1.tar.gz.

File metadata

  • Download URL: redditflow-0.0.1.tar.gz
  • Upload date:
  • Size: 9.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.0 CPython/3.8.13

File hashes

Hashes for redditflow-0.0.1.tar.gz
Algorithm Hash digest
SHA256 e2ef212bb6e6ae14462b771270a5a34bfa9e9b408019302bcbf7511ab9433bce
MD5 c4043097410393bd0f030970a09d386a
BLAKE2b-256 20ec41ba4e055aea200ee3594b974debd9ad3576458339f2d0404b003f4ecb6c

See more details on using hashes here.

File details

Details for the file redditflow-0.0.1-py3.8.egg.

File metadata

  • Download URL: redditflow-0.0.1-py3.8.egg
  • Upload date:
  • Size: 28.8 kB
  • Tags: Egg
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.0 CPython/3.8.13

File hashes

Hashes for redditflow-0.0.1-py3.8.egg
Algorithm Hash digest
SHA256 14a5c411132c7b791fd792db1ac45162fcf5282c36cb75faa681674d0d45e76e
MD5 b009f4cfa1ce8b8cd5beebcd82306ff7
BLAKE2b-256 2ed6e76edfd4f2a3a2ea200788b257c383e750a0e08ded8e3ee8fd5383092ed9

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page