Skip to main content

Quora-scraper

N|Solid

Quora-scraper is a command-line application written in Python that scrapes Quora. It simulates a browser environment to let you scrape Quora rich textual data. You can use one of the three scraping modules to: Find questions that discuss about certain topics (such as Finance, Politics, Tesla or Donald-Trump). Scrape Quora answers related to certain questions, or scrape users profile. Please use it responsibly !

Install

To use our scraper, please follow the steps below:

$ pip install quora-scraper

To update quora-scraper:

$ pip install quora-scraper --upgrade

Alternatively, you can clone the project and run the following command to install: Make sure you cd into the quora-scraper folder before performing the command below.

$  python setup.py install

Usage

quora-scraper has three scraping modules : questions ,answers,users.

1) Scraping questions URL:

You can scrape questions related to certain topics using questions command. This module takes as an input a list of topic keywords. Output is a questions_URL file containing the topic's question links.

Scraping a topic questions can be done as follows:

  • a) Use -l parameter + topic keywords list.

    $ quora-scraper questions -l [finance,politics,Donald-Trump]
    
  • b) Use -f parameter + topic keywords file location. (keywords must be line separated inside the file):

    $ quora-scraper questions -f  topics_file.txt
    

2) Scraping answers:

Quora answers are scraped using answers command. This module takes as an input a list of Questions URL. Output is a file of scraped answers (answers.txt). An answer consists of :

Quest-ID | AnswerDate | AnswerAuthor-ID | Quest-tags | Answer-Text

To scrape answers, use one of the following methods:

  • a) Use -l parameter + question URLs list.

    $ quora-scraper answers -l [https://www.quora.com/Is-milk-good,https://www.quora.com/Was-Einstein-a-fake-and-a-plagiarist]
    
  • b) Use -f parameter + question URLs file location:

    $ quora-scraper answers -f  questions_url.txt
    

3) Scraping Quora user profile:

You can scrape Quora Users profile using users command. The users module takes as an input a list of Quora user IDs. The output is UserProfile file containing:

First line : UserID | ProfileDescription |ProfileBio | Location | TotalViews |NBAnswers | NBQuestions | NBFollowers | NBFollowing

Remaining lines (User's answers): AnswerDate | QuestionID | AnswerText

Scraping Users profile can be done as follows:

  • a) Use -l parameter + User-IDs list.

    $ quora-scraper users -l [Albert-Einstein-195,Jackie-Chan-8]
    
  • b) Use -f parameter + User-IDs file.

    $ quora-scraper users -f quora_username_file.txt
    

Notes

a) Input files must be line separated.

b) Output files fields are tab separated.

c) You can add a list/line index parameter In order to start the scraping from that index. The code below will start scraping from "physics" keyword: sh $ quora-scraper questions -l [finance,politics,tech,physics,life,sports] -i 3

d) Quora website puts limit on the number of questions accessible on a topic page. Thus, even if a topic has a large number of questions (ex: 100k), the number scraped questions links will not exceed 2k or 3k questions.

e) For more help use :

   $ quora-scraper --help

f) Quora-scraper uses xpaths and bs4 methods to scrape Quora webpage elements. Since Quora HTML Structure is constantly changing, the code may need modification from time to time. Please feel free to update and contribute to the source-code in order to keep the scraper up-to-date.

License

This project uses the following license: MIT

Metadata

Release files for quora-scraper 1.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for quora-scraper 1.1.3
File Size Uploaded
quora-scraper-1.1.3.tar.gz 5.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for quora-scraper 1.1.3
File Interpreter ABI Platform
quora_scraper-1.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 10.5 MB

Release files / quora-scraper-1.1.3.tar.gz

Download URL quora-scraper-1.1.3.tar.gz
Size 5.2 MB
Tags Source
SHA-256 checksum
How to use checksums
aa0bafed1604cbbc70b40b0ac0d9f068b3303a5f1beca5e0985e7252f36e10a0
BLAKE2b-256 checksum
How to use checksums
2fad10819ca9ca8e4b9676fa7ac15ff5ae03d1493f5983793ec6f780b05cc99a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.2.0 pkginfo/1.5.0.1 requests/2.24.0 setuptools/46.1.3 requests-toolbelt/0.9.1 tqdm/4.47.0 CPython/3.6.5

Release files / quora_scraper-1.1.3-py3-none-any.whl

Download URL quora_scraper-1.1.3-py3-none-any.whl
Size 5.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
539a7b20b1819b09d2a299bd965491f855878081574b1b55332de49edbd2b583
BLAKE2b-256 checksum
How to use checksums
8f7184aecb9e3a3cc73772f5c1a1b12085b96b6742a8ff0abfbeb6e21ca5c95d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.2.0 pkginfo/1.5.0.1 requests/2.24.0 setuptools/46.1.3 requests-toolbelt/0.9.1 tqdm/4.47.0 CPython/3.6.5

Release history Release notifications | RSS feed

This release

1.1.3 This release

2 release files

1.1.2

2 release files

1.1.0

2 release files

1.0.8

2 release files

1.0.6

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page